Assemble a narrated-UGC "stitch reply" ad from a config — a single spoken VO carries a verbatim testimonial while ~30 per-cut i2v clips (one creator across ~5 wardrobes in ~3 worlds, plus product B-roll) are each trimmed to their EDL window built from the VO's Whisper word boundaries and hard-concatenated via filter_complex concat (never the demuxer, which drops audio on a duration mismatch), the VO mixed over an optional sidechain-ducked instrumental bed (−20dB, 20 to 1) so the VO stays on top, karaoke-pop captions burned on every word throughout (VEED Whisper preset, re-spelled against the locked script), a landing-page scroll rendered as FFmpeg zoompan over a Playwright PNG (not i2v), and closed on the brand's real end-card PNG — never AI-rendered text. This is the FREE deterministic assembly stage (trim-to-EDL + filter_complex concat + VO and music mix + karaoke captions + landing-page zoompan + end-card append); the VO, creator, start-frames, and clips come from create-vo-elevenlabs / create-image-gpt-image-fal / create-image-fal / create-video-fal. Use for the narrated-ugc-wardrobe-stitch format.
Assemble a narrated-UGC "stitch reply" ad from a config: a fast-cut vertical testimonial where a single spoken VO carries a verbatim ~13-sentence reversal-hook monologue over ONE creator across ~5 wardrobe changes in ~3 micro-worlds, interspersed with product B-roll (capsule macro, unboxing, a landing-page scroll), ~30 hard cuts on the VO cadence, closing on a brand end card. This capability is the FREE, deterministic assembly — trim-to-EDL, hard-concat, the VO+music mix, the karaoke-pop caption burn, the landing-page zoompan, and the end-card append.
scripts/config.example.json is the worked example (Bioma "Do NOT buy Bioma Probiotics", ~37s
1080×1920 9:16, ~30 body cuts + a ~2s end card); scripts/PIPELINE.md maps every config block to
its source step and scripts/README.md documents the free assembly.
This is the FREE, deterministic assembly stage — it spends nothing. The paid inputs are
separate capabilities — the spoken VO (create-vo-elevenlabs) Whisper-aligned so the WORD
BOUNDARIES set the cut grid; one locked creator (create-image-gpt-image-fal anchor + ~5 wardrobe
edits chained off the anchor) + 3 world wides + per-cut start-frames (create-image-fal product
composites); and one Veo/Seedance i2v clip per cut (create-video-fal). Given the VO +
vo-final.words.json + edl.json + one clip per cut + a Playwright landing-page PNG + the brand
end-card PNG, render-narrated-ugc-wardrobe-stitch trims each clip to its EDL window, hard-concats
on the VO cadence, mixes the VO over the ducked bed, burns the karaoke-pop captions, appends the
end card → the master. Re-cuts reuse the existing VO / start-frames / clips and cost $0.
hook, feature,
reaction-insert, payoff-hold, b-roll-insert, landing-page); snap every cut window to the
word boundaries. The payoff line gets a HELD payoff-hold beat (~3× mean shot length).filter_complex concat, not the demuxer. Trim each clip to its EDL window and
hard-concat with filter_complex concat — the -f concat demuxer drops the audio when a
drawtext/scale step shaves a clip a few ms below its window. No dissolves.vo-final.words.json (VEED
Whisper preset, bold yellow), on every word; re-spell brand tokens Whisper mishears against the
locked script ("synbiotic" over "symbiotic"; keep "I'ma" verbatim) — never edit the script to
match Whisper. Captions are suppressed over the end card. If VEED mis-captions a brand token,
hand-patch that sentence with local ASS karaoke.filter_complex concat, VO+music mix,
caption burn, landing-page zoompan, end-card append, loudnorm I=-14 → a 1080×1920 h264+aac
master (~37s). No paid calls, no keys.8866b2a
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.