Mandatory pre-publish review gate for a UGC video render. Transcribes the finished render's AUDIO with Whisper and word-diffs it against the approved spoken script, then gates pinning the final render (video_project_upsert patch.final_render_id) — blocking a render whose generated audio mis-voices a word (e.g. the approved "human-vetted" spoken as "human witted"), says a different number or brand name, flips a negation, drops an approved phrase, or comes back silent. Correct speech written differently ("5mg" said "five milligrams", "30%" said "thirty percent", a spoken URL, "don't" said "do not", "braxleybands" said "braxley bands", a confirmed pronunciation like "AG1" said "A G one") passes. Runnable, gating counterpart to content-goose's review-transcript-integrity atom. Every ugc-video-formats recipe runs this after render and BEFORE pinning the final render.
70
88%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
Medium
Suggest reviewing before use
The QC gate every UGC video recipe MUST clear before it publishes. Not an eyeball
/watch— a deterministic transcript-vs-script diff that exits non-zero on a defect so the recipe can hard-stop pinning the final render (video_project_upsertpatch.final_render_id).
In short: it compares what the render actually SAYS with the script the user approved. Correct speech that is only written differently passes (numbers, units, URLs, contractions, fused brand names, confirmed pronunciations). A wrong number, a flipped negation, a mis-voiced word or brand, a dropped phrase or silence fails.
Seedance generates the audio natively. It sometimes mis-voices a word — the
approved line human-vetted comes back spoken as human witted; documented
siblings: Hume→Hune, Alitu→al-too. The defect lives in the render's
audio, so an eyeball /watch ("dialogue matches the script") slips it through,
and a downstream caption pass then bakes the wrong word in verbatim. Nothing was
comparing the actual spoken audio against the script the user approved.
This gate does exactly that, deterministically, and refuses to publish on a miss.
remix-ugc-*-from-sample and create-ugc-*-video-from-refs
recipe, in the QC phase, after the master render exists and before
pinning it as the final render (video_project_upsert patch.final_render_id).Before rendering, persist the exact approved spoken lines (the verbatim utterance,
no beat notes) to working/approved-script.txt. Then, after render:
python3 <pack>/review-ugc-render/scripts/review_render.py \
--video working/final.mp4 \
--script-file working/approved-script.txt \
--json working/review-verdict.jsonWhen the voice-over used confirmed pronunciations, add them (recommended):
--pronunciations working/brand-rules.json — the same file create-vo-elevenlabs
read (its --rules brand-rules.json, or the read_pronunciations.py output; both
carry a pronunciations: [{term, say_as}] list). Leave the flag out when there is
no such file: a missing file is an ERROR (exit 3).
video_project_upsert patch.final_render_id).Transcription backend (in priority order): the GooseWorks whisper-proxy (CLI
credentials or the sandbox token) → OPENAI_API_KEY (honors OPENAI_BASE_URL) →
local whisper CLI. ffmpeg must be on PATH.
Both the script and the transcript are put in one canonical spoken form before the diff. The rules are bounded — each is an exact rewrite, never a fuzzy match.
| Written | Heard | Rule |
|---|---|---|
49, 105, 2,500, 1 million | forty-nine, one hundred and five, two thousand five hundred, a million | number words = digits |
2.5, 2026, 249, 1st, 2nd | two point five, twenty twenty six, two forty-nine, first, second | decimals, years and prices read in pairs (only when spoken as words: written 2 20-minute is never 220), ordinals (second = 2nd only when the script writes 2nd) |
No. 1, No.1, #1 | number one | number sign, only when written with the dot or # (no 1-star reviews and no one stay negations) |
5mg, 30g, 500ml, 12oz, 10 lbs, 30-day | five milligrams, thirty grams, …, thirty days | unit after a quantity (weights, volumes, %, money, hours/minutes/seconds, days/weeks/months/years, calories, x times) |
30% | thirty percent / 30 per cent | percent |
$49, $49.99 | forty nine dollars, forty nine dollars and ninety nine cents | money |
braxleybands.com, www.example.com | braxleybands dot com, w w w dot example dot com, example dot com | URL; www. is optional |
don't, can't, it's, you're | do not, cannot / can not, it is, you are | contractions |
Braxleybands, Gooseworks, everyone, OneSkin | Braxley Bands, goose works, every one, One Skin | fused/split: exact join of 2–3 words |
AG1 | AG1, AG one, A.G. one | written forms, no alias needed |
AG1 | A G one, A G 1 (letters spaced out) | only with a confirmed alias |
Guards that keep the rules honest:
| Report line | Root cause | Fix |
|---|---|---|
[high] said "59" where script has "49" — number differs… | Wrong, added or dropped number | Re-roll. A number is never a benign paraphrase. |
[high] dropped "5mg" — unit differs… | A unit after a number was changed, added or dropped ("5mg" said "five") | Re-roll. Only a dropped dollars/euros/pounds alone ("$9.99" said "nine ninety-nine") is not HIGH; dropped cents is. |
[high] extra "doesn't" … negation changed | A not/never/no/without was added or lost — the claim flips | Re-roll. |
[high] said "Hune" where script has "Hume" — brand name not heard as approved | Brand mis-voiced or dropped (--brand-term / confirmed pronunciation) | Re-roll; spell it phonetically in the SPOKEN LINE (e.g. Ali-too, never a (pronounced …) parenthetical). See create-video-seedance-2-fal Failure Modes. |
[high] said "witted" where script has "vetted" — audio likely mis-voices… | Seedance mis-voiced a similar-looking word | Re-roll a new seed. |
[medium] dropped "…" / low similarity | Seedance dropped an approved phrase | Re-roll; if only a tail word, a surgical stitch_replacement.py window fix may recover it. |
[low] extra "…" + low similarity | Extra speech beyond benign filler | Re-roll. A single filler word ("so", "okay") in a normal-length line is LOW and passes. |
⚠ audio is effectively silent | Wrong render / audio track lost in post | Re-render / re-check the mux; never publish a silent take. |
ERROR: no transcription backend | No proxy credentials, no OPENAI_API_KEY, no local whisper | Sign in / set the key / install whisper, then re-run. |
ERROR: alias … would change a number, unit or negation | A bad --alias or pronunciation entry | Fix the alias. Aliases may only respell a name. |
The verdict passes only when similarity ≥ --min-ratio and there is no HIGH issue.
Numbers, units, negations and declared brand names fail on their own (HIGH). Any
other dropped or extra word is medium/low: it lowers the similarity, and fails the
gate only when the similarity drops below --min-ratio (so one dropped ordinary
word in a long line can pass).
--expect-music is advisory only; it does not by itself fail the gate.
Brand words are never removed from the diff. (Before 2026-10-06, --brand-term
fuzzily stripped brand-like words, which let "Hune" pass for "Hume". That is gone.)
--pronunciations PATH (recommended when the voice-over used one) — the
confirmed pronunciations file: create-vo-elevenlabs's brand-rules.json or the
output of its scripts/read_pronunciations.py
({"brand_id", "basis", "pronunciations": [{"term", "say_as", "fact_id"}]}).
Every entry becomes an alias and its term a brand term. An entry that cannot be used
(see alias rules) is skipped with a WARNING, never fatal. A missing or unreadable
file is an ERROR (exit 3), so pass the flag only when the file exists.--alias "TERM=SPOKEN" (repeatable) — one confirmed spoken form, e.g.
--alias "AG1=A G one". The term also becomes a brand term. A bad --alias is an
ERROR (exit 3), checked before any transcription is spent.--brand-term TERM (repeatable) — marks a brand name. Exactly:
[low], counted as a
match) only when all of these hold: it has no negation, number or unit
change; every word on both sides is an alphabetic word of some
--brand-term; and the two sides look alike (character similarity ≥ 0.6). This
keeps the older calling pattern working: a script written in the spoken form
(Try ak-mee today) with --brand-term Acme --brand-term ak --brand-term mee
passes when Whisper writes Try Acme today ("ak mee" vs "acme" = 0.67).
Everything else still fails HIGH: a heard word that is not a declared term
(Hume heard Hune), two different declared names (Hims vs Hers, Hume Body Pod vs Hume Band), a truncation (Acme heard ak), or a number inside a term
(Pod 4 vs Pod 5, 7-Eleven: 7 days vs 11 days).
Prefer --pronunciations over passing say_as words as --brand-term.Rules for aliases:
not, never, without, none, nothing
or nobody ("Hume" = "never Hume"; no and nor are fine as syllables, "Nomad" =
"no-mad"), or when the spoken form is only numbers, units or negations ("Decagon" =
"five").--video PATH (required) — the rendered master mp4.--script-file PATH or --script "text" — the approved spoken script. Omit
both only for a genuinely script-free clip (the drift check is then skipped and
the gate is advisory).--brand-term TERM, --alias "TERM=SPOKEN", --pronunciations PATH — see above.--min-ratio FLOAT (default 0.90) — transcript↔script similarity to pass,
measured on the canonical spoken form.--captions-srt PATH — optional SRT to check for caption-text defects.--json PATH — write the machine verdict for the app's review panel. Each issue
carries the canonical tokens (script_words / heard_words) and the original
wording (script_text / heard_text).python3 tests/test_review_render.py # or: python3 -m pytest tests/Pure Python, no audio, network or paid call. Covers the QA-71 audit fixture table (units, URL, numbers, percent, contractions, fused brands, wrong price, negation, Hume→Hune), number words, units (including added/dropped units in long lines), URLs, negation flips, fused/split words, CLI-style brand terms, confirmed AG1 aliases and saved pronunciations with number-like syllables, omission, extra speech, the report wording, and the CLI exit codes with transcription stubbed out.
This is the shipped, single-file, gating slice of the fuller
coworkers/video/molecules/review/review-loop (18-axis rubric). Here we enforce
the one axis that catches audio-vs-script defects at publish time
(review-transcript-integrity / brand_text_accuracy). Deeper multi-axis review
stays in the content-goose lab.
cc3e518
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.