CtrlK
BlogDocsLog inGet started
Tessl Logo

whisper

OpenAI's general-purpose speech recognition model. Supports 99 languages, transcription, translation to English, and language identification. Six model sizes from tiny (39M params) to large (1550M params). Use for speech-to-text, podcast transcription, or multilingual audio processing. Best for robust, multilingual ASR.

56

Quality

66%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./backend/cli/skills/llm-tools/whisper/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

53%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A solid, example-driven reference skill with genuinely executable code for installation, transcription, CLI usage, and integration, but it carries marketing-style padding, an inlined language list that shadows an unreferenced bundle file, and a batch-processing workflow with no validation steps. The biggest structural fix is wiring references/languages.md into the body and trimming metrics/performance content Claude does not need.

Suggestions

Replace the inline 'Language support' list with a pointer: 'Full list with WER tiers: see [languages.md](references/languages.md)' — the file exists in the bundle but is never referenced.

Add validation to the Batch processing section (guard against missing files, check the transcription succeeded, confirm each output file was written) so the batch workflow is not unguarded.

Cut the Metrics section (GitHub stars, training hours), the real-time-factor performance table, and the duplicate language list — this is reference/marketing material Claude does not need to transcribe audio.

DimensionReasoningScore

Conciseness

The body is mostly dense, executable code, but includes padded sections Claude does not need — 'Metrics' ('72,900+ GitHub stars', 'Trained on 680,000 hours of audio'), a real-time-factor performance table, and an inline 11-language list that duplicates the bundled references/languages.md. This fits 'mostly efficient but includes some unnecessary explanation or could be tightened' rather than anchor 2, since the core content is code rather than explanatory prose.

3 / 5

Actionability

Nearly all examples are copy-paste executable ('pip install -U openai-whisper', 'whisper audio.mp3 --output_format srt', 'ffmpeg -i video.mp4 -vn -acodec pcm_s16le audio.wav') and cover the common cases. The LangChain integration example imports 'WhisperTranscriptionLoader' from 'langchain.document_loaders', which is not a standard, verifiable class — a minor gap that keeps this below anchor 5.

4 / 5

Workflow Clarity

As a reference-style skill the sections are well organized, but the Batch processing section loops over files with no validation — no missing-file check, no error handling, no confirmation that output was written. Per the rubric, batch operations without validation cap workflow clarity at 3, which takes precedence over the simple-skill exception.

3 / 5

Progressive Disclosure

The bundle contains references/languages.md (189 lines), yet the body never links to it — the 'Language support' section inlines a partial list and says only 'Full list: 99 languages total'. The reference is effectively buried and its content is inlined in the body, matching 'content that clearly belongs in separate files is inlined; or references are buried' rather than anchor 3, which requires at least some signaled reference.

2 / 5

Total

12

/

20

Passed

Description

78%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description that clearly states what the skill does and when to use it, with concrete, natural trigger terms and a distinct ASR niche. The main improvement is sharpening the 'when' clause into explicit user-trigger phrasing and adding a few common synonyms (transcribe audio, subtitles, .mp3) to broaden keyword coverage.

Suggestions

Rephrase the trigger clause to user-trigger form: 'Use when the user asks to transcribe audio, subtitles, or podcasts, or mentions speech-to-text, .mp3/.wav files, or multilingual ASR.'

Add common synonyms users actually say — 'transcribe', 'audio file', 'meeting notes', 'captions/subtitles' — to broaden trigger coverage toward anchor 5.

DimensionReasoningScore

Specificity

It names the domain ('OpenAI's general-purpose speech recognition model') and lists several concrete capabilities — 'transcription, translation to English, and language identification', 'Six model sizes from tiny (39M params) to large (1550M params)'. Minor gaps in coverage (no CLI, output formats, or timestamps) keep it at anchor 4 rather than 5, and it clearly exceeds the 1-2 actions of anchor 3.

4 / 5

Completeness

Both parts are present: the 'what' is explicit ('Supports 99 languages, transcription, translation to English, and language identification') and the 'when' is stated ('Use for speech-to-text, podcast transcription, or multilingual audio processing'). The 'when' is capability-oriented rather than user-trigger phrasing ('Use when the user asks to transcribe...'), so it fits anchor 4 — both present but 'when' could be more explicit — rather than anchor 5.

4 / 5

Trigger Term Quality

Good natural keyword coverage — 'speech-to-text, podcast transcription, or multilingual audio processing', 'multilingual ASR' — phrased the way users actually ask. A few natural terms are missing ('transcribe audio', 'audio file', '.mp3', 'subtitles', 'captions', 'meeting notes'), which matches anchor 4 rather than the comprehensive synonym-and-extension coverage of anchor 5.

4 / 5

Distinctiveness Conflict Risk

'Best for robust, multilingual ASR' plus triggers like 'speech-to-text' and 'podcast transcription' carve out a clear niche that only adjacent audio/ASR skills could overlap, matching anchor 5 ('clear niche with distinct triggers; minimal conflict risk'). It is more distinctive than anchor 4's 'minor overlap risk with closely related skills' examples.

5 / 5

Total

17

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

14

/

16

Passed

Repository
synthetic-sciences/openscience
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.