CtrlK
BlogDocsLog inGet started
Tessl Logo

whisper

Transcribe and translate speech in 99 languages.

52

Quality

59%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./optional-skills/mlops/whisper/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

65%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A strong, example-driven reference: almost every section is executable code or a concrete command sequence, and it stays lean by mostly avoiding background explanation. The main weaknesses are the orphaned references/languages.md file that the body never points to (while duplicating its content), the absence of any validation or error-recovery step in the batch-processing workflow, and one dubious integration example (LangChain's WhisperTranscriptionLoader).

Suggestions

Replace the inline top-10 language list with a signaled one-level reference, e.g. '**Full language list & per-language WER**: See [languages.md](references/languages.md)', so the bundle file is actually reachable and duplication is removed.

Add error handling to the batch-processing example (try/except per file, checking result['text'] is non-empty before writing) so the batch workflow has validation checkpoints and can score above the cap of 3.

Verify or remove the LangChain example — 'WhisperTranscriptionLoader' does not appear to be a real LangChain class; replace it with a working documented pattern or drop the section.

DimensionReasoningScore

Conciseness

The body is dominated by tight, executable code blocks and tables with minimal prose ('# Requires ffmpeg', model-size table, CLI flag examples), and it largely assumes Claude's competence rather than explaining what Whisper or ASR is. It is not 5 because of padding that earns no tokens — the 'Metrics' block ('72,900+ GitHub stars', 'Trained on 680,000 hours of audio'), a duplicated Resources list, and a top-10 language list that repeats what lives in references/languages.md; not 3 because the waste is minor relative to the whole.

4 / 5

Actionability

Nearly all guidance is copy-paste executable: install commands, load_model/transcribe snippets with result-structure access, per-option examples, full CLI invocations, and a batch loop — matching 'Mostly executable guidance; concrete code or commands with minor gaps'. Not 5 because of gaps: the LangChain example uses 'WhisperTranscriptionLoader' which is not a real documented class, and the batch example has no error handling around individual files.

4 / 5

Workflow Clarity

The single-action core flow (install → load model → transcribe) is unambiguous, but the skill includes batch operations, and the batch-processing section has no validation or error-recovery checkpoint — per the guidelines, missing validation in batch workflows caps workflow clarity at 3. It also stays at 3 rather than 4 because multi-step flows like 'extract audio from video → transcribe → generate subtitles' are scattered across sections with no explicit sequence or checkpoints.

3 / 5

Progressive Disclosure

A bundle file exists (references/languages.md) and the body's Language support section duplicates its content (a top-10 list plus 'Full list: 99 languages total') without ever linking or signaling the reference file — matching 'references present but not clearly signaled; content that should be separate is inline'. Not 4 because an available, well-structured reference is left unnavigable from the SKILL.md; not 2 because the body itself is well-sectioned and the reference is only one level deep.

3 / 5

Total

14

/

20

Passed

Description

53%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is concise and names two genuine, concrete capabilities, but it lacks any 'Use when' trigger guidance and omits common natural phrasings (transcription, speech-to-text, subtitles, audio formats) that would help it fire on real user requests. It is serviceable but below the standard of the good reference examples.

Suggestions

Add an explicit trigger clause, e.g. 'Use when the user wants to transcribe audio or video files, generate subtitles, or translate speech to English.'

Include natural synonyms and formats users actually say — 'transcription', 'speech-to-text', 'subtitles', 'podcast/meeting audio', '.mp3/.wav/.mp4' — to broaden trigger coverage.

Mention one or two additional concrete capabilities already documented in the body (e.g. word-level timestamps, SRT/VTT subtitle output) to raise specificity.

DimensionReasoningScore

Specificity

The description states 'Transcribe and translate speech in 99 languages' — it names the domain (speech) and two concrete actions (transcribe, translate), matching the anchor 'Names domain and 1-2 concrete actions, but not comprehensive'. It does not reach 4 because other real capabilities shown in the body (word timestamps, SRT/VTT subtitle generation, batch processing, language detection) are absent; it stays above 2 because the two actions it does list are concrete, not generic.

3 / 5

Completeness

It has a clear 'what' ('Transcribe and translate speech in 99 languages') but no 'Use when...' clause or equivalent explicit trigger guidance, which the judging guidelines state caps completeness at 3. It is not 4 because the 'when' is entirely absent rather than merely under-specified.

3 / 5

Trigger Term Quality

'Transcribe', 'translate', 'speech', and '99 languages' are relevant natural keywords, but common variations users would actually say are missing — 'transcription', 'speech-to-text', 'audio files', 'subtitles', and file extensions like .mp3/.wav — matching the anchor 'Some relevant keywords but missing common variations or synonyms'. Not 4, since the omission is more than 'a few natural terms'; not 2, since the terms present are domain-specific rather than generic.

3 / 5

Distinctiveness Conflict Risk

Speech transcription/translation is a fairly clear niche with distinct triggers ('transcribe', 'speech', 'translate... to English'), giving only minor overlap risk with closely related audio/speech skills — matching 'Mostly distinct; minor overlap risk'. Not 5 because the description does not include trigger phrases specific enough to fully disambiguate from other ASR or translation skills.

4 / 5

Total

13

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

13

/

16

Passed

Repository
NousResearch/hermes-agent
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.