Assemble a narrated motion-graphic LISTICLE video ad from a config — a spoken voiceover (cloned or cast, in the chosen tone) carries a numbered listicle while N web-animated hyperframe beats (HTML plus the Web Animations API, one branded design system of alternating tiles, big hero numerals, and callouts in the chosen visual style) are rendered frame-by-frame via Playwright and anchored to the VO's word-level timestamps, optional color-graded breather windows (brand clip, AI clip, or motion-graphic) give visual breath, and captions burn ONLY inside those B-roll windows (2-word chunks, ASS Format header carrying a Name field so none drop) with the VO mixed under a low music bed. This is the FREE deterministic assembly stage (Playwright beat render plus ffmpeg concat plus window-masked caption burn plus VO-and-music mix plus final composite) — the VO, the music bed, and any AI breather clips come from create-vo-elevenlabs, create-music-elevenlabs, and create-video-fal. Use for the vo-anchored-motion-listicle format.
Assemble a narrated motion-graphic listicle ad from a config: a spoken voiceover carries a numbered listicle (hook + N points + CTA) and every visual beat is anchored to
the VO's word-level timestamps. Each beat is a web-animated hyperframe (an HTML page + the Web
Animations API driven by window.renderAt(t)) rendered to video frame-by-frame with Playwright, all
in ONE branded design system (alternating background tiles, big hero numerals, body type, decorative
accents, callouts). Optional color-graded breather windows give visual breath, and
captions burn only on the B-roll windows. The shipped master is pure motion-graphic + VO — there
is NO lipsync (the still headshot is kept only for a future lipsync variant). This capability
is the FREE, deterministic assembly — the Playwright beat render, the ffmpeg concat, the
window-masked caption burn, the VO+music mix, and the final composite.
scripts/config.example.json is the worked example (Everself "doctor-educator" listicle, ~66s
1080×1920 9:16 at 25fps) — copy its structure, never its creative values; scripts/PIPELINE.md maps every config block to its source step and
scripts/README.md documents the free assembly.
The creative calls are made upstream (the format recipe's choices, asked of the user) and arrive in
the config; this assembly just renders what it is given. The demo's values are examples only.
voice.mode / voice.voice_id.
The demo used a cloned doctor-educator.beats[].
The demo used expert tips with mistake → fix beats.design.name / callout_style /
accents / background_tiles. The demo used EDUCATOR_MOTION with glass-pill callouts.music.prompt (none = mix the VO alone). The demo used a warm lo-fi
bed at ~0.18.This is the FREE, deterministic assembly stage — it spends nothing. The paid inputs are separate
capabilities — the spoken VO (create-vo-elevenlabs, a cloned or cast voice, eleven_v3 +
atempo) whose word-level timestamps (Groq whisper-large-v3 word-level) set the timeline; the low
music bed (create-music-elevenlabs); and any AI breather clips (create-video-fal, trimmed +
color-graded; brand clips and motion-graphic breathers are free). There is no free stock-footage
source.
Given the VO + words-flat.json + the N authored hyperframe beats + the color-graded B-roll windows +
the brand wordmark SVG, render-vo-anchored-motion-listicle renders each beat frame-by-frame via
Playwright (all beats at fps 25), concats the beats + B-roll, burns the window-masked captions, mixes
the VO under the low music bed, and composites → the master. Re-cuts reuse the existing VO / beats /
B-roll and cost $0.
window.renderAt(t); Playwright screenshots it frame-by-frame and
ffmpeg encodes it. This is NOT i2v — it is deterministic web motion graphics._shared.css (palette + type + alternating
tiles + accents + callout style) so N beats read as one designed reel; alternate only the background
tile, keep numerals / body / accents / callouts consistent.Format: header MUST carry a Name field — without it the leading-comma bug eats the
first field and captions silently drop. If the host ffmpeg lacks libass, render the cues as timed
PIL PNG overlays (ffmpeg overlay=…:enable='between(t,st,en)') at the same placement.none) sits ~0.18 vol under it. No ducking needed at that level.I=-14 → a 1080×1920 25fps h264 crf18 + aac 192k master. No paid calls, no keys.cc3e518
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.