Bilal Labs / Videos with Claude Opus 5.5
Captioned vertical clip with local TTS voiceover
A short script becomes a vertical clip for X, Reels or Shorts: a local text-to-speech voice, word timings from local Whisper, captions that light up word by word, and a code card that types in. Nothing leaves the machine.
Transcript: I made this video without opening a video editor. I gave Claude Opus 5.5 one prompt. It wrote HTML, and HyperFrames rendered it to MP4. The captions, the voice, the motion. All local. Here is the exact recipe.
The prompt
Make a 15-second 1080×1920 vertical clip with HyperFrames from this script: "I made this video without opening a video editor. I gave Claude Opus 5.5 one prompt. It wrote HTML, and HyperFrames rendered it to MP4. The captions, the voice, the motion. All local. Here is the exact recipe." Generate the voiceover locally, get word timestamps with npx hyperframes transcribe, and show big centered captions in the lower third, 3–6 words per group, highlighting each word as it is spoken. Above them, a code card whose lines fade in. Dark background, one warm accent. Run check, then render.
This is the prompt to reproduce it. I built it inside a longer multi-task agent session, so there is no per-recipe token count to report.
What the agent did
- Generated the voiceover with macOS say (voice Samantha, rate 175) and converted it to WAV with FFmpeg. HyperFrames' own Kokoro TTS needs a Python package that would not install on the system Python here, so the agent used the built-in voice.
- Ran npx hyperframes transcribe (whisper.cpp, small.en) for word timestamps. Whisper heard "5.5 one" as "5.51", so that word was fixed by hand.
- Wrote index.html: caption groups built from the word list, each word switching to full white at its start time.
- Placed the audio as an <audio> clip with an id (without one the render is silent).
- Ran npx hyperframes check (84/84 contrast checks passed), trimmed the audio slot to the 14.76 s voice length, rendered.
Measured numbers
- Render time: 14.5 s on an M4 MacBook (hardware GPU)
- Output: 1080×1920, 30 fps, H.264 + AAC, 15.0 s, 1.6 MB
- Transcription: whisper.cpp small.en, local, 37 words
- Tokens / cost: Not measured: built inside a larger agent session with no per-recipe counter
Notes and pitfalls
- Check the transcript before you trust it. Numbers and product names are where Whisper slips.
- Every <audio> needs an id or the HyperFrames mixer skips it.
- Keep captions in the middle-lower area. The bottom ~15% sits under platform UI on Reels and Shorts.
Source
Prompt, source and notes: opus-video-recipes/captioned-vertical-clip. Framework choice: HyperFrames vs Remotion vs Manim.
FAQ
Can I use a better voice?
Yes. Swap assets/vo.wav for any recording or TTS output, rerun npx hyperframes transcribe on it, and paste the new word timings into the W array.
Are the captions burned in?
Yes, they are part of the video. The transcript is also on this page as text.
Does this need an API key?
Not for rendering. The voice, the transcription and the render all run locally. You only need a model to write the HTML.