hyperframes-media
Asset preprocessing for HyperFrames compositions — multi-provider TTS (HeyGen / ElevenLabs / Kokoro local), multi-provider BGM (Google Lyria / local MusicGen), Whisper transcription, background removal, and caption authoring. Use for npx hyperframes tts, bgm, transcribe, remove-background, voice/provider selection, music-mood prompting, captions / subtitles / lyrics / karaoke / per-word styling.
다음 행동
/hyperframes-media기술 README 원문 보기
설치 옵션, 예시 코드, 세부 사용법을 영어 README 원문 그대로 확인합니다.
HyperFrames Media
CLI commands that create assets (tts, bgm, transcribe, remove-background), plus everything needed to consume and animate transcript data in HTML. For placing assets into compositions, see hyperframes-core.
Provider chains (auto-detected from env)
TTS — npx hyperframes tts "..." picks the first available provider:
| Order | Provider | Detected when | Word timestamps |
|---|---|---|---|
| 1 | HeyGen (Starfish) | $HEYGEN_API_KEY / hyperframes auth login | Yes, native — pass --words narration.words.json to capture |
| 2 | ElevenLabs | $ELEVENLABS_API_KEY set | No — chain transcribe after |
| 3 | Kokoro-82M (local, 54 voices) | always (no key required) | No — chain transcribe after |
> If the installed hyperframes tts is the local-only build (its --help says "Kokoro-82M" and has no --provider/--words flags), it silently falls back to Kokoro even with $HEYGEN_API_KEY set. To force HeyGen regardless of CLI version, use the self-contained scripts/heygen-tts.mjs (see references/tts.md).
BGM — npx hyperframes bgm --duration N:
| Order | Provider | Detected when |
|---|---|---|
| 1 | Google Lyria (RealTime) | $GEMINI_API_KEY or $GOOGLE_API_KEY set |
| 2 | MusicGen (facebook/musicgen-small, local) | Python transformers + torch + soundfile installed |
Override either with --provider <name>.
Routing
| Task | Read |
|---|---|
npx hyperframes tts — provider chain, voice IDs, words.json | references/tts.md |
| HeyGen without the CLI — self-contained REST script (wav + words) | scripts/heygen-tts.mjs (see references/tts.md) |
npx hyperframes bgm — Lyria vs MusicGen, mood prompts, tuning | references/bgm.md |
npx hyperframes transcribe — Whisper, model rules, output shape | references/transcribe.md |
npx hyperframes remove-background — transparent cutouts | references/remove-background.md |
| TTS → transcription → captions (no recorded voiceover) | references/tts-to-captions.md |
| Caption authoring — style detection, layout, word grouping, exit | references/captions/authoring.md |
| Transcript handling — input formats, quality gates, cleanup, APIs | references/captions/transcript-handling.md |
| Caption motion — karaoke, marker effects, audio-reactive | references/captions/motion.md |
| Model caches, system dependencies, troubleshooting | references/requirements.md |
Non-negotiable rules
- Voice IDs are provider-specific.
am_michaelis Kokoro-only; HeyGen UUIDs don't work on Kokoro. If you pass--voice, also pin--providerto avoid silent provider drift when the user's env changes. - Always pass `--model` to `transcribe`. The CLI default
small.ensilently translates non-English audio. Seereferences/transcribe.md→ "Language Rule". - HeyGen returns word timestamps; ElevenLabs / Kokoro do not. When you want captions, either pass
--wordsto HeyGen and use that JSON directly, or runtranscribeagainst the audio file. Don't assume word data is always there. - Captions consume the flat word-array format with
{ id, text, start, end }. Seereferences/transcribe.md→ "Output Shape". - `remove-background --background-output` is hole-cut, not inpainted. For "scene without the person", a different tool is needed. See
references/remove-background.md→ "When NOT the right tool".
이것도 같이 보면 좋다
같은 업무 태그와 카테고리가 겹치는 항목부터 보여줍니다.

agents-sdk
Build AI agents on Cloudflare Workers using the Agents SDK. Load when creating stateful agents, durable workflows, real-time WebSocket apps, scheduled tasks, MCP servers, chat applications, voice agents, or browser automation. Covers Agent class, state management, callable RPC, Workflows, durable execution, queues, retries, observability, and React hooks. Biases towards retrieval from Cloudflare docs over pre-trained knowledge.

faceless-explainer
faceless-explainer video workflow - arbitrary text (article / notes / topic / brief) -> narrator_scripts.json + audio (voice + BGM) + section_plan.md -> typography / abstract-graphics / diagram / data-viz video. Typical length up to ~3 min (sweet spot ~30-90s); a genuinely longer piece is general-video, not this workflow. Generates its OWN narration (TTS) — it does not sync to a user-supplied / pre-recorded voiceover (that is general-video). No website capture, no real product screenshots. If the text names a product / its site to promote, that is /product-launch-video; when product-vs-topic is unclear, start at /hyperframes-read-first.

hyperframes-creative
Non-animation creative direction for HyperFrames videos. Use for design spec (frame.md / design.md) handling, palettes, typography, narration, beat planning, audio-reactive visuals, composition patterns, and brand / style decisions. For atomic motion patterns and scene blueprints, use `hyperframes-animation`.