CtrlK
BlogDocsLog inGet started
Tessl Logo

caption-burn

Burned-in captions for a finished vertical video, three kinds. transcribe.py gets word timings from the video's own audio through the GooseWorks proxy (fal Whisper, bills the Ads agent, cents); captions.py burns one to three words at a time with Pillow + ffmpeg (no libass needed), either pinned to a split-screen seam (plate 25% above / 75% below) or at a fixed height, in a plate, outline or one-word serif style, with an optional red hook card; plates.py burns per-beat caption blocks (black, one union silhouette, placed in the emptiest band) for formats with no voice. The last caption (the CTA) holds to the final frame. Use as the last step of any video ad.

75

Quality

94%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

96%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An exemplarily lean, command-first skill body: every run command is executable, defaults and constraints are quantified, and the batch-render workflow includes explicit validation and feedback loops. The only structural weakness is that the footprint helper is not named or invocable from the page, unlike the other bundled scripts.

Suggestions

Name the footprint script explicitly (scripts/footprint.py) and show its command line in a code block, as is done for transcribe.py/captions.py/plates.py, since 'the bundled footprint helper' is currently uninvocable from the page.

Mention scripts/media_proxy.py and scripts/_fonts.py in one line (e.g. what media_proxy.py serves and when _fonts.py fetches Roboto) so every bundle file is discoverable from SKILL.md.

DimensionReasoningScore

Conciseness

The body is dense with operational detail and zero padding: no concept explanations, no library intros, every sentence carries a flag, default, or constraint (e.g. "plate 25% above / 75% below", "cap ~0.019 H", "~14 characters a line"). It matches the score-5 anchor — lean, assumes Claude's competence, every token earns its place — and is well above the score-4 'minor instances of over-explanation' case.

5 / 5

Actionability

Every workflow is given as a copy-paste-ready command with real flags and defaults: `python captions.py --video reel.mp4 --beats cutlist.aligned.json --words reel.words.json --out final.mp4 [--style plate|outline]`, plus concrete card syntax `--card "LINE ONE|LINE TWO" --card-until 4.7`. The beats input format is specified ("Any file with `beats: [{start, end, vo}]` works"), covering the common cases like the description's examples do at score 5.

5 / 5

Workflow Clarity

The sequence (transcribe finished audio → captions burn → inspect) is explicit, and the batch render operation has real validation checkpoints and feedback loops: "regenerate after alignment or copy changes", "Preview rejects a changed cut list until its footprint is rebuilt", "Inspect final captioned frames as well as this planned coverage", and "An explicitly supplied missing font fails in both commands" documents error behavior. This matches the score-5 anchor (explicit validation steps plus error-recovery loops) for a batch operation that would otherwise be capped at 3.

5 / 5

Progressive Disclosure

The body is well organized into Run / Placement and style / plates.py / Rules / Footprint sections and the scripts it names (transcribe.py, captions.py, plates.py) exist in the scripts/ bundle with commands shown. However the footprint tool is referred to only as "the bundled footprint helper" without naming footprint.py or giving its command line, and media_proxy.py/_fonts.py are never mentioned — minor organization gaps that fit the score-4 anchor rather than the fully clear score-5 navigation.

4 / 5

Total

19

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An unusually specific, third-person description that names every script, style, placement mode and billing implication, and closes with an explicit 'use as the last step of any video ad' trigger. The only weakness is keyword coverage: the common synonyms 'subtitles'/'CC' and video file extensions are absent.

Suggestions

Add the synonyms users most often say for this task — e.g. 'subtitles' and 'CC' — so the skill triggers when someone asks to 'burn subtitles' rather than 'captions'.

Mention the video file extensions handled (e.g. .mp4/.mov) in the description, since trigger-term coverage currently stops at the generic word 'video'.

DimensionReasoningScore

Specificity

The description lists multiple specific concrete actions with comprehensive coverage: "gets word timings from the video's own audio", "burns one to three words at a time with Pillow + ffmpeg (no libass needed)", "pinned to a split-screen seam (plate 25% above / 75% below) or at a fixed height", "plate, outline or one-word serif style", "optional red hook card", and "burns per-beat caption blocks (black, one union silhouette, placed in the emptiest band)". It matches the score-5 anchor of multiple specific concrete actions covering all three caption kinds plus the transcription step.

5 / 5

Completeness

It clearly answers what (three caption kinds, which script does what, styles, placement, CTA hold behavior) and when via an explicit trigger clause: "Use as the last step of any video ad." The 'when' is explicit and concrete rather than weakly implied, matching the score-5 anchor; without that clause completeness would have been capped at 3.

5 / 5

Trigger Term Quality

Natural terms a user would say are present — "captions", "burned-in captions", "vertical video", "video ad", "transcribe", "no voice" — but common variations are missing: no "subtitles" or "CC" synonyms, and no file extension (".mp4", ".mov"). This fits the score-4 anchor (good keyword coverage, a few natural terms missing) rather than 5 (comprehensive synonyms and extensions) or 3 (missing common variations broadly).

4 / 5

Distinctiveness Conflict Risk

It occupies a clear niche — burned-in captions for finished vertical video ads via named scripts (transcribe.py/captions.py/plates.py) with a GooseWorks/fal Whisper proxy and per-beat caption blocks — with triggers unlikely to fire for unrelated skills. It matches the score-5 anchor (clear niche, distinct triggers, minimal conflict risk); it is far more specific than the score-4 'Works with PDF and Word document files' example.

5 / 5

Total

19

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
gooseworks-ai/goose-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.