Assemble a split-screen creator ad from a config — a two-zone vertical composite where a supplied AI-creator lip-sync take fills the BOTTOM ~48% while real 16:9 product/demo clips run uncropped in the TOP ~52%, each top clip contain-fit with a darkened blurred cover-scale fill of the same clip (never black bars), a 3px brand-color divider between the zones, the creator slice cover-fit per the per-scene VO timing, scenes hard-concatenated with the body audio being the concatenated creator VO slices, an end card held on the last sharp frame ~3s, then the ASSEMBLED cut transcribed with local Whisper (not the raw VO — concat drops inter-scene silence) and word-level captions burned in the chosen style. This is the FREE deterministic assembly + caption stage (two-zone composite + blurred fill + divider + hard-concat + end card + captions); the VO comes from create-vo-elevenlabs, the anchor from create-image-gpt-image-fal, and the whole-VO lip-sync from a paid VEED Fabric 1.0 take (a no-atom upstream input). Use for the split-screen-creator format.
Assemble a split-screen creator ad from a config: a two-zone vertical (1080×1920, 9:16, ~40s) format where an AI creator talking-head anchors the BOTTOM ~48% of the frame and real 16:9 product/demo clips run uncropped in the TOP ~52%, a 3px brand-color divider between the zones. The creator delivers the whole VO cold-to-camera and each top clip proves the claim its VO line makes. This capability is the FREE, deterministic assembly + captions — the two-zone composite (contain-fit + blurred-cover fill + divider + creator slice), the hard-concat, the end card, and the word-level caption burn from the assembled cut.
scripts/config.example.json is the worked example (Perplexity concept-10
"Bloomberg terminal", ~40s 1080×1920 9:16, 6 scenes + an end card);
scripts/PIPELINE.md maps every config block to its source step and
scripts/README.md documents the free assembly.
This is the FREE, deterministic assembly + captions stage — it spends
nothing. The paid inputs are separate steps — the VO (create-vo-elevenlabs,
ElevenLabs eleven_v3 with-timestamps, sliced into per-scene windows), the
photoreal MEDIUM chest-up AI-creator anchor (create-image-gpt-image-fal,
gpt-image-2) — shot at a natural webcam distance (headroom + shoulders, real
room), not a plain-background close-up headshot (see the anchor note below), and
the whole-VO lip-sync (a paid VEED Fabric 1.0 @ 720p take — a no-atom step,
image_url = the anchor, audio_url = the vo mp3, ~$0.15/sec, ~$5.90 for a 39s
VO; run its calls sequentially, veed/fabric-1.0 storage-auths 403 under
parallel load). Given the creator lip-sync take + the per-scene VO timing + one
16:9 top clip per scene + the scene-1 hook graphic + the end-card clip,
render-split-screen-creator composites the two zones, hard-concats the scenes,
appends the end card, transcribes the assembled cut, and burns the captions → the
master. Re-cuts reuse the existing VO / lip-sync / clips and cost $0.
top_height ~998. Keep every stacked height EVEN
(998 + 4 divider + 918 = 1920) — libx264 rejects odd dimensions.ugc-walk-and-talk). VEED Fabric handles photoreal fine (unlike Seedance).top_start/top_end) to the on-message segment that proves its VO line.
Never loop a short clip — set the window and the assembler speed-fits it to
the scene (looping replays into a sparse/black tail).timing.json; the
lip-sync drives the mouth.endcard.clip_end),
not the black tail.serif-accent, kinetic-pop, …). Keep the -precaption cut + the
.ass sidecar so captions restyle without re-rendering the composite. If the
host ffmpeg lacks libass, render the cues as timed PIL PNG overlays (ffmpeg
overlay=…:enable='between(t,st,en)') at the same placement.loudnorm I=-14 → a
1080×1920 h264+aac master. No paid calls, no keys — the VEED Fabric lip-sync is
a supplied input, produced upstream.8866b2a
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.