Comprehensive patterns for AI-powered audio generation including text-to-music, voice synthesis, text-to-speech, sound effects, and audio manipulation using MusicGen, Bark, ElevenLabs, and more. Use when "music generation, text to music, AI music, voice cloning, text to speech, TTS API, ElevenLabs, MusicGen, Bark, audio synthesis, sound effects generation, voice synthesis, AudioCraft, " mentioned.
How to make an analog-horror PSA short — the stenciled-pictogram / robo-broadcast / VHS-overlay format (faceless, 10 scenes × ~3s, 9:16, ~30s total). A domain overlay on the standard pipeline that supplies the IF / DO-NOT / BUT / AND scenario structure, the locked yellow-on-black 1970s-civil-defense pictogram visual STYLE, the chroma-key-to-alpha rule (icons must be transparent PNGs before composition), the robo-PSA voice profile (ElevenLabs "Alerter" community voice + ALL CAPS input + stability ~0.5 + style 0), the layered VHS-noise overlay stack (SnowCanvas + VcrTrackingCanvas + MobiusScanlines + MobiusWobble), the 5-layer RGB chromatic-split for icons + captions, the SMPTE colour-bars climax (3 wide vertical bars + filter blur(2.5px) mandatory), and the `-tune grain` CRF 30 final-encode rule for noise-heavy renders. Works for ANY PSA topic — "your fridge is not your fridge", "your phone is not your phone". USE WHEN the user asks for an analog horror PSA / "creepy PSA-style short" / fake-emergency-broadcast video / VHS-aesthetic warning / stenciled-pictogram horror short for TikTok / Reels / Shorts, with no specific existing video to reproduce. This is a niche SKILL (generalized), not a remix TEMPLATE. For "remix this exact PSA but swap the subject", use the remix path in docs/skills-vs-templates.md.
Update or polish the AOSP Internals reference book (the chapters at the repo root -- NN-slug.md -- plus the appendices and the generated agents/ Part-skills) so it is accurate and current for a newly-synced or newer AOSP/Android source tree. Use this whenever the user says anything like "polish the book for Android NN", "update the reference book to the new AOSP version", "the AOSP source got re-synced, bring the book up to date", "deep-rewrite the chapters against the new source", or -- importantly -- "cover the new Android modules/projects in the book". It drives a resumable, changeset-prioritized, per-chapter convergent deep-rewrite with an agent-team accuracy + Mermaid review loop, and MANDATORILY runs a gap analysis that finds newly-added modules/projects and folds them into existing chapters (or adds new chapters/Parts only by an explicit threshold). Trigger it even when the user only says "update the book for the new Android" without naming this process, and even when they only want the new-modules part.
Audio-first long-form explainer pipeline — takes one audio file (or a YouTube/podcast URL → audio) and produces a faceless overlay-driven video in the dev-essay / tech-podcast style (5-30 minutes, 16:9, no talking head). Orchestrates the existing `ralphy` primitives: yt-dlp pull → silence-remove → ElevenLabs Scribe v1 word-level transcript → Gemini audio-describe → LLM claim segmentation → per-claim overlay-type planner (code-block / terminal / tweet-card / browser-frame / screenshot / meme / diagram / quote-card / chapter-card / logo-pop) → asset prep (Playwright for screenshots, `ralphy generate image` for memes, ElevenLabs Music for the bed, ElevenLabs SFX for whoosh/pop/hit) → HyperFrames composition with chapter sub-compositions → `ralphy render`. USE WHEN the user drops an audio file (mp3 / wav / m4a) or a YouTube / podcast URL and asks for a long-form video edited on top, says "make a video from this audio / podcast", "monter ce podcast" (FR), "smontiruy podkast v video" (RU transliteration), "audio to video", "long-form essay video", "faceless overlay video from audio". TRIGGER (EN): "audio to video", "podcast to video", "edit my podcast into a video", "make a long-form video from this audio", "faceless explainer from audio", "overlay-driven video from podcast", "monter le podcast", "smontiruy podkast". See body for ALSO FIRE / DO NOT FIRE / HARD INVARIANTS.
Cinematic B-roll shot planner for NanoBanana. Analyzes an uploaded image (STYLE_ANCHOR) or user text to produce exactly 5 cohesive, edit-ready B-roll shot outputs.
This skill provides the five-pillar language system for constructing video prompts with cinematic precision. Apply it to any video generation task — atmosphere shots, product reveals, character scenes, abstract motion, sacred/spiritual visuals, or any brief where imprecise language would produce unpredictable results.
The master workflow for Jamey "Quix". Transforms CoD screenshots into 3D Blender-style renders, composites them onto backgrounds, and applies a heavy YouTube thumbnail enhancement stack focused on weapon sharpness and vibrant (but controlled) environments.
DyNote: systematically and efficiently extract raw Douyin/DY video data and analyze videos, comments, accounts, hashtags, and short-video scenes into evidence-graded learning notes, summaries, research briefs, scripts, and knowledge-base material. Use when the user asks to 抓取/提取/整理 抖音视频字幕、视频文案、ASR 转写、Qwen3-ASR 中文转写、原始材料归档、学习笔记、analysis plan、note budget、避免返工、复用已有素材、评论洞察、账号分析、赛道/话题研究、竞品拆解、电商/本地生活视频分析、事实核查、自动搜索素材, save Douyin content as Markdown/TXT, or use subtitle/local ASR as the factual spine with logged-in Douyin Web built-in AI / Doubao fallback as visual or quick-reading supplements.
Quality evaluation of rendered UGC mp4s — scene segmentation, audio loudness / dead-air, caption density, and per-scene visual analysis. Produces an actionable report (eval.json + eval-report.md) sized for a downstream fixer agent. USE WHEN the user asks to "evaluate / score / grade / review / QA / check quality of" a rendered video, asks "is this video good?", drops an mp4 path with no other instruction, mentions "find issues / problems / artifacts", asks for retention or scroll-stop assessment, or has just rendered something and wants verification before publishing. TRIGGER (EN): "evaluate this video", "score the render", "grade the mp4", "review the final cut", "QA this video", "is this ready to ship", "what's wrong with this video", "find issues in <path.mp4>", "audit the video", "scene-by-scene breakdown", "retention check", "quality gate". See body for ALSO FIRE / DO NOT FIRE / HARD INVARIANTS.