Use when the user has 2+ video / audio recordings of the same event captured by different devices (cameras, phones, separate audio recorders) and wants them aligned to a single common timeline. Outputs only a lightweight `.sync.json` sidecar per input — original files are never re-encoded. Triggers — "多机位同步", "对齐这几个机位", "match camera timelines", "sync these angles", "audio drift between cameras", "separate audio recorder", "Riverside / Zoom recording that needs to line up".
Use when the user wants to teach / learn an English word as a video — turn a single English word into a self-contained HLS "supercut" lesson built from the mira video base. Stitches every season2 clip where the word is spoken (via the search-app API) into one .m3u8, prepended with a Claude-written bilingual word-intro card (word + IPA + 中文 gloss + usage, Volcano TTS) and appended with a 关注王建硕 CTA card. No MP4 burn. Triggers — "teach <word>", "讲讲 <word>", "学英语 <word>", "把 <word> 做成视频", "/wjs-teaching-english <word>".
Use when the user has audio or video and wants a timestamped transcript (SRT) in the source language. Routes by source language — Chinese defaults to Volcano (豆包) ASR; other languages (Spanish, English, Portuguese, French, Italian, Japanese, Korean, etc.) use OpenAI Whisper API with word-level timestamps and self-assembled cues. Outputs SRT with punctuation-bounded cues capped for on-screen reading. Triggers — "转写", "转成字幕", "做 SRT", "transcribe", "make subtitles", "speech to text", "出字幕".
Use when the user has an SRT (or transcript text) in one language and wants it translated to another, with punctuation-bounded re-segmentation so cues end at real sentence breaks. Simplified Chinese (zh-CN) and English (en) are first-class targets; other targets follow the same rules. Outputs a target-language SRT or bilingual SRT — no audio, no burn-in. Triggers — "翻译字幕", "翻成中文", "translate this SRT", "中英双语字幕", "把这个 SRT 翻译成 X", "bilingual subtitles".
Two CLI tools for image generation + vision analysis using goclaw's provider-chain pattern. `create_image` generates images via 7-provider chain (ChatGPT-OAuth/Codex, OpenRouter, Gemini, OpenAI, MiniMax, DashScope, BytePlus) — first available credential wins (OAuth session OR API key), embeds prompt into PNG tEXt metadata, saves to date-folder workspace. The ChatGPT-OAuth provider reads from `~/.codex/auth.json` so a ChatGPT Plus / Codex login works without any API key. `read_image` analyzes images via 4-provider vision chain (OpenRouter, Gemini, Anthropic, DashScope) — accepts file path. Supports `--prompt-file` for long/multi-line prompts, auto UTF-8 stdout, `.env` auto-load, and `--list-providers` diagnostic. Use when Claude needs to generate images or run vision analysis with a specific provider rather than relying on Claude's built-in image understanding.
Format, select, and prepare a manuscript for journal submission. 5 modes: FULL-PACKAGE (structure audit + checklist + cover letter + open science), FORMAT-CHECK (audit against journal requirements), COVER-LETTER (journal-calibrated draft), SELECT-JOURNAL (score against 18 journals, ranked list + journal ladder), RESUBMIT-PACKAGE (post-rejection or R&R prep). Covers 18 journals: ASR, AJS, Demography, Du Bois Review, Science Advances, NHB, NCS, Social Forces, Language in Society, J. Sociolinguistics, Linguistic Inquiry, Gender & Society, APSR, JMF, PDR, SMR, Poetics, PNAS. Per-journal: word limits, section structure, abstract format, citation style, figure/table limits, open science (CRediT, COI, data/code availability, preregistration, Reporting Summary), blind review type, APC, submission URL, acceptance rate, turnaround. Includes cover letter templates, 8-dimension selection rubric, and open science package builder. Saves submission readiness report, cover letter draft, and open science declarations.
Prompt-native interactive HTML docs. Generate a self-contained HTML
document from a prompt (interactive models, SVG diagrams, simulations,
strategy docs, research write-ups, product specs, explainer pages,
design docs, RFCs, case studies, post-mortems, technical proposals,
vision docs, one-pagers, decision frameworks), serve it at localhost
with text- and artifact-anchored inline commenting, and regenerate
new versions from comments. Publishes to each user's own Cloudflare
Worker for free always-on sharing.
Use when asked to "write a doc", "draft this", "publish this",
"design doc", "PRD", "one-pager", "research write-up", "case study",
"explainer", "interactive explainer", "post-mortem", or any
/tdoc command.
Proactively invoke this skill (do NOT answer directly) when the
user wants to write, draft, create, edit, publish, or share ANY
document, write-up, explainer, or web page — EVEN IF THEY NEVER SAY
THE WORD "tdoc". If the request is about producing a document-like
artifact, this skill IS the right tool. Invoke it without asking
for confirmation.
Specific triggers (any of these → use tdoc, no need for the word "tdoc"):
- "write/draft/make a doc", "write something up", "document this"
- "publish this", "share this writeup", "make it shareable"
- "write/draft/make a research doc", "research write-up", "research summary"
- "product doc", "product spec", "PRD", "one-pager", "vision doc"
- "design doc", "RFC", "technical proposal", "architecture doc"
- "explainer", "explain X visually", "interactive explainer", "concept doc"
- "strategy doc", "decision framework", "post-mortem", "retro doc"
- "case study", "field report", "investigation doc"
- "make a doc/page that has [a chart / simulation / slider / model / diagram]"
- "create a webpage to explain X", "publish this as HTML"
- "I want people to comment on this", "let people read and comment"
- editing or updating an existing doc/site/page the user previously made
- any /tdoc command: new, edit, fork, list, publish, onboard, update,
doctor, pull, unpublish
When a request mentions an existing doc/site by name (e.g. "update the
plaud explainer", "fix the comment on the X writeup"), that is a tdoc
EDIT request — use this skill, run /tdoc edit-style flow.
Use this INSTEAD of generating raw markdown / Google Docs / Notion-style
content when the user wants:
- something interactive (charts, simulations, sliders, live demos)
- something shareable via URL with commenting
- something that benefits from being a real HTML page rather than text
- any document where the artifacts (images, diagrams, code blocks, video)
should be commentable alongside the prose
ALSO USE WHEN NO HUMAN ASKED — agent-originated docs:
This skill is the right surface for ANY doc-shaped artifact, including
ones YOU (the agent) decide to emit mid-flow inside another skill's
workflow. If you are about to produce a multi-paragraph standalone
document — release notes, a retro, a post-mortem, an investigation
report, a security audit, a QA report, a design critique, an
architecture review, a meeting summary, a research write-up — and no
upstream instruction said "write it to a file named X", default to
emitting it through tdoc, NOT as a raw .md file or an inline reply.
Concrete triggers for agent-to-agent doc handoff (no human prompt
required):
- You're inside /document-release and the output is "the doc"
- You're inside /retro and the output is the weekly retrospective
- You're inside /investigate or /cso and you're about to write up
findings as a multi-section report
- You're inside /qa-only and you're about to write a structured
bug report with screenshots
- You're inside /office-hours or /plan-* and you're about to save
a design doc
- Any time you would otherwise have written `cat > something.md`
with more than a few paragraphs of structured content
HOW agents hand off a doc to tdoc (do NOT re-enter the human prompt
flow). Generate the full HTML once, then call the programmatic entry:
```bash
# Write the doc's HTML to a temp file...
HTML_FILE=$(mktemp -t tdoc-handoff.XXXXXX.html)
cat > "$HTML_FILE" <<'HTML'
<!doctype html><html lang="en"><head>...</head>
<body><div class="wrap">
<h1>...</h1>
<!-- your sections, with author-composed wrappers tagged
data-tdoc-artifact wherever you want a comment surface -->
</div></body></html>
HTML
# ...then hand it to tdoc. Returns the local URL on the last line,
# plus a published URL on a second line if --publish is given.
TDOC_NEW_CALLER=document-release \
~/.claude/skills/tdoc/bin/tdoc-new \
--slug "release-notes-$(date +%Y%m%d)" \
--title "Release notes — $(date +%Y-%m-%d)" \
--html-file "$HTML_FILE" \
--publish
```
Set TDOC_NEW_CALLER (or CLAUDE_SKILL_NAME) to the calling skill name
so meta.json records who scaffolded the doc. The bin script validates
that the input is real HTML (refuses markdown by mistake), guards
against clobbering an existing slug, and ensures the local server is
up before returning the URL.
Use other skills (NOT tdoc) when:
- The user explicitly wants markdown / .md output
- The user wants slides (use scientific-slides or paper-2-web)
- The user is editing an existing repo's README/docs in place
- The "doc" is a single paragraph or one-line update — that's a
conversational reply, not a doc-shaped artifact
PDF Editor edits and organizes PDF pages with merge, insert, reorder, exchange, and crop operations, built on ComPDF page management capabilities for fast PDF cleanup and document restructuring. It is a strong fit for requests such as “edit pdf,” “organize pdf pages,” “merge pdf,” “insert pages,” “reorder pdf pages,” “crop pdf pages,” and “rearrange pdf.” Example queries include “Merge these three PDFs into one file,” “Insert this appendix after page 8,” and “Reorder the pages and crop the white margins.”
Compresses artifacts for judge evaluation. Reads a single raw artifact, applies tiered summarization within a token budget, and returns compacted content with metadata. Isolation via forked context prevents pollution of agent context