CtrlK
BlogDocsLog inGet started
Tessl Logo

create-vo-elevenlabs

Generate a voiceover (VO) clip via ElevenLabs text-to-speech, ROUTED THROUGH THE elevenlabs-proxy so it bills the Ads agent. Voice id + script text come from the template recipe. Use for the spoken narration of VO-driven video-ad formats (cgi-app-sizzle, flat-vector-explainer, hypermotion). Never call ElevenLabs directly — the proxy attribution is required.

65

Quality

82%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

75%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A tight, domain-dense body with genuine validation discipline around the paid call (fresh brand reads, persistence confirmation, dry run, error on missing alignment). Its main defects are execution gaps in the timestamp and pronunciation-export features: the timestamp flag is never named, read_pronunciations.py has no invocation example, and the referenced 'bundled timestamp tests' do not exist in the bundle.

Suggestions

Replace 'Use the timestamp option' with the actual flag and a copy-paste example, e.g. gen_vo.py --text "..." --voice <id> --out vo.mp3 --with-timestamps (plus --settings for per-voice settings), and state that the alignment lands in <out-stem>.timestamps.json.

Remove or fix the 'See the bundled timestamp tests' reference — no test file ships in the bundle; instead inline the request/output contract (the alignment JSON's keys and the same-stem output file) in one or two lines.

Add a one-line invocation for the bundled read_pronunciations.py (its --context, --brand, --require and --out arguments) so the pronunciation-export step is as executable as the gen_vo.py steps.

DimensionReasoningScore

Conciseness

The 31-line body is lean and assumes competence: executable commands, no explanation of what TTS or ElevenLabs is, and every sentence carries non-obvious operational policy (proxy billing, phonetic-spelling rule, persistence re-read, dry run before spending). Not score 4 because there is no padded or over-explained passage to trim — apparent repetition (the captions-keep-written-name rule) is deliberate emphasis in two different contexts.

5 / 5

Actionability

The primary and pronunciation flows are fully executable — 'gen_vo.py --text "Meet Drinkag1." --voice <id> --out vo.mp3 --rules working/brand-rules.json' and '--say-as "Drinkag1=drink A G one"' (both verified against the bundled script) — but key details are missing elsewhere: the timestamp feature is described only as 'Use the timestamp option' without naming the actual flag (--with-timestamps) or giving an invocation, no command line is shown for the bundled read_pronunciations.py, and 'See the bundled timestamp tests for the request and output contract' points to a file that does not exist in the bundle. This lands at 'concrete guidance but incomplete / missing key details' rather than score 4's 'minor gaps'.

3 / 5

Workflow Clarity

The risky paid-call workflow is well sequenced with explicit checkpoints: read fresh brand facts at the start ('A previous video's notes or a cached rules file do not establish current pronunciation'), resolve before 'the paid voice step', 'Save the user's confirmed choice... then make a separate read and confirm it survived', a free dry run to review spoken text, and 'A missing alignment is an error; do not render captions from guessed timings'. Not score 5 because the sequence lives in prose sections rather than an explicit ordered flow with error-recovery loops; validation is present, so the missing-validation cap at 3 does not apply.

4 / 5

Progressive Disclosure

A sub-50-line skill with no external reference files needed, organized into three well-labeled sections, and the bundled scripts it names (gen_vo.py, read_pronunciations.py, media_proxy) all exist. It falls short of 5 because one reference is dangling: 'See the bundled timestamp tests' — no test file ships in scripts/ (only gen_vo.py, media_proxy.py, read_pronunciations.py), so that navigation target fails against the actual bundle structure.

4 / 5

Total

16

/

20

Passed

Description

82%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: concrete action, explicit 'Use for...' trigger with named video-ad formats, and a hard routing boundary that makes it highly distinct. The only weakness is capability breadth — it lists just the core generate action rather than the fuller feature set (pronunciation rules, timestamps) the skill actually supports.

DimensionReasoningScore

Specificity

Names the domain and a concrete action — 'Generate a voiceover (VO) clip via ElevenLabs text-to-speech, ROUTED THROUGH THE elevenlabs-proxy' — plus the sourcing detail 'Voice id + script text come from the template recipe', but only 1-2 concrete actions are listed, not a comprehensive set (no mention of timestamped output, pronunciation handling, or dry runs). It is above score 2 because the action is specific and instrumented (proxy routing for Ads-agent billing), not generic, and below score 4 because it does not enumerate several specific capabilities.

3 / 5

Completeness

Explicitly answers both: what — 'Generate a voiceover (VO) clip via ElevenLabs text-to-speech' routed through the proxy — and when — 'Use for the spoken narration of VO-driven video-ad formats (cgi-app-sizzle, flat-vector-explainer, hypermotion)', a concrete 'Use for...' trigger clause. It also adds an explicit boundary ('Never call ElevenLabs directly'). This matches the score 5 anchor with concrete trigger phrases, not merely the weaker 'when' of score 4.

5 / 5

Trigger Term Quality

Good natural keyword coverage: 'voiceover', 'VO clip', 'ElevenLabs', 'text-to-speech', 'spoken narration', 'video-ad formats', plus named formats (cgi-app-sizzle, flat-vector-explainer, hypermotion). A few natural terms a user might say are missing, e.g. 'TTS', 'voice-over', 'narration audio', so it falls short of the comprehensive synonym/extension coverage of score 5.

4 / 5

Distinctiveness Conflict Risk

Clear niche — VO generation for specific video-ad formats, exclusively via the elevenlabs-proxy with required attribution — with distinct named triggers. The explicit 'Never call ElevenLabs directly' boundary further separates it from any generic TTS skill, giving minimal conflict risk per the score 5 anchor.

5 / 5

Total

17

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
gooseworks-ai/goose-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.