Assemble a multi-scene GRWM beauty-demo ad from a config — a locked-identity creator applies ~5 products step by step while a SEPARATE ElevenLabs voiceover narrates and every scene cut is snapped to the VO's product-name word-starts (Whisper word-level timestamps), then ~5 Playwright product overlay cards (real PDP-verified taglines) are composited onto the master each on its product-NAME word-start, the SEPARATE VO is mixed on top of a ducked music bed at loudnorm I=-14, clean-white 3-words/cue captions are burned, and the video closes on a flat-lay end card. This is the FREE deterministic assembly stage (re-cut to the VO word-starts, hard-concat, Playwright card render + card composite, VO plus music mix, caption burn, flat-lay end card); the VO, scene clips, product cutouts, and music come from create-music-elevenlabs / create-image-gpt-image-fal / create-video-fal. Use for the glassy-matte-grwm format.
64
77%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
View guide
Low
Low-risk findings worth noting
Fix and improve this skill with Tessl
tessl review fix ./skills/ads/capabilities/render-glassy-matte-grwm/SKILL.mdAssemble a multi-scene GRWM beauty-demo ad from a config — a locked-identity creator applies ~5 makeup/skincare products step by step at a vanity, a separate ElevenLabs voiceover narrates the routine, and every scene cut is snapped to the VO's product-name word-starts, with ~5 Playwright product overlay cards on the product-name beats, a ducked music bed, burned captions, and a flat-lay end card. This capability is the FREE, deterministic assembly — the Whisper-driven re-cut + hard-concat, the Playwright card render + card composite, the VO + music mix, the caption burn, and the flat-lay end card.
This is the multi-scene beauty demo, distinct from the single-take apparel outfit-reveal
(ugc-grwm, one Seedance reference-to-video call with native lip-sync and minimal post). Here the
timeline is driven by a SEPARATE VO and the scenes are re-cut to its word-starts.
scripts/config.example.json is the worked example (DIBS Beauty "5-Step Glassy Matte Routine",
~32s 1080×1920 9:16, 12 VO-snapped cuts + 5 product cards); scripts/PIPELINE.md maps every
config block to its source step and scripts/README.md documents the free assembly.
This is the FREE, deterministic assembly stage — it spends nothing. The paid inputs are
separate capabilities — the SEPARATE narration VO (create-music-elevenlabs, or a user-supplied
mp3; word-level Whisper timestamps set the timeline), ~7 Seedance scene clips one per product step
(create-video-fal), the ~5 white-bg product cutouts + the flat-lay end-card still
(create-image-gpt-image-fal), and the ducked music bed. Given the VO + .words.json + one clip
per step + the ~5 product cutouts + the music bed, render-glassy-matte-grwm re-cuts each clip to
its VO word-start window, hard-concats on the cut, renders + composites the product cards on the
product-name beats, mixes the VO over the ducked music, burns the captions, and appends the
flat-lay end card → the master. Re-cuts reuse the existing VO / clips / cutouts and cost $0.
-c:v libx264 -crf 20 — -c copy corrupts the duration when zoompan/PNG clips are in the
chain.-loop 1 -t <dur> — without it the PNG emits one frame at t=0 and the fade/enable filters
silently no-op (cards go invisible).loudnorm I=-14. If the host ffmpeg lacks a filter, apad/atrim to length before the mix.overlay=…:enable='between(t,st,en)' at the same
placement.8866b2a
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.