Assemble a song-driven music-video ad from a config — a generated sung track carries the whole narration across N tableaux (one keyframe -> one i2v clip per lyric beat) with NO separate voiceover, captions synced to the song's OWN word timings (script-window, never Whisper) and the hook word landing on the chorus drop, closed on a PIL brand end card. This is the FREE deterministic assembly stage (clip cut-to-timeline + captions + end card + FFmpeg composite); the song, keyframes, and clips come from create-music-elevenlabs / create-image-fal / create-video-fal. Use for the song-driven-music-video format.
60
75%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
Passed
No findings from the security scan
Fix and improve this skill with Tessl
tessl review fix ./skills/ads/capabilities/render-song-mv/SKILL.mdAssemble a song-driven music-video ad from a config: a purpose-written, sung song is
the entire narration (no separate voiceover), and every visual beat is timed to the lyrics.
The delivered song sets the timeline; N tableaux (one keyframe → one image-to-video clip per
lyric beat, all in a single look pack) are cut to their lyric windows and hard-concatenated
on the beat, captions are built from the song's OWN word timings with the hook line landing
on the chorus drop, and the spot closes on a PIL brand end card. It reads like a tiny animated
music video, not a demo. scripts/config.example.json is one worked example (Loóna "Fall In
Love With Sleep Again", 28s paper-craft 9:16) — its style, song, vocalist, mood, protagonist and
setting are that demo's answers, not defaults; scripts/PIPELINE.md maps every config block
to its step and scripts/README.md documents the free assembly.
The creative calls are made upstream by the user (the recipe's choices) and arrive in the config;
this assembly never picks them.
look_pack, clip_engine.motion_opener. The demo used a paper-craft diorama
at night.song.prompt / song.bpm. The demo used a dreamy synth lullaby at ~80 BPM.song.prompt. The demo used a soft breathy female lead.song.prompt, the lyrics, tableaux. The demo went calm/dreamy.look_pack.style_opener. The demo used a young-woman paper-doll.tableaux[].keyframe_prompt. The demo went bedroom at night → moonlit paper village.This is the FREE, deterministic assembly stage — it spends nothing. The three paid
inputs are separate capabilities: the sung song (create-music-elevenlabs, music_v1,
force_instrumental FALSE — the lyrics ARE the script, returns mp3 + words.json), one
keyframe per tableau (create-image-fal), and one Kling 3.0 i2v clip per tableau
(create-video-fal). Given the delivered song + words.json + one clip per beat,
render-song-mv cuts each clip to its lyric window, hard-concats on the beat, builds the
lyric-synced captions, composites the PIL end card, and muxes → the master. Re-cuts reuse
the existing song / keyframes / clips and cost $0.
force_instrumental false); do not add a spoken voiceover or a
second music bed.timeline.json) — never trim the song to a pre-planned grid.audio/words.json (~3 words at lyric boundaries); accent words get captions.accent_color.
Whisper on sung audio returns "🎵 Music Playing 🎵", so it can't caption lyrics.is_hook) is timed so the
payoff word (song.hook_word) sits on the chorus drop; accent that word in the captions.style_opener + negative_tail + palette drives
every keyframe so N beats read as one film; no morph within a clip.audio_mix.climax_beat_id — set it to
this run's hook tableau; the demo's T08 is example-only), loudnorm to −14 LUFS → 1080×1920
h264+aac. No paid calls, no keys.cc3e518
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.