CtrlK
BlogDocsLog inGet started
Tessl Logo

openai-whisper

Local speech-to-text with the Whisper CLI (no API key).

60

Quality

73%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/openai-whisper/SKILL.md

The canonical home for this skill is openai-whisper in openclaw/openclaw

SKILL.md
Quality
Evals
Security

Quality

Content

92%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is an exemplary lean CLI quick-start: two fully executable commands covering the main use cases, followed by short non-obvious notes (cache location, model default). The only blemishes are a generic model-size platitude and an install-specific default that could be trimmed.

DimensionReasoningScore

Conciseness

The body is lean with executable commands up front, but "Use smaller models for speed, larger for accuracy" is knowledge Claude already has, and "--model defaults to turbo on this install" is version-sensitive detail — minor instances of over-explanation that could be trimmed. Not 5 for those few tokens that don't earn their place; well above the noticeable padding of anchor 3.

4 / 5

Actionability

The two Quick start commands ("whisper /path/audio.mp3 --model medium --output_format txt --output_dir ." and "whisper /path/audio.m4a --task translate --output_format srt") are fully executable, copy-paste ready, and cover the common cases of transcription and translation with concrete flags. This matches the anchor-5 example structure of complete commands covering common cases.

5 / 5

Workflow Clarity

This is a simple, single-purpose skill and the single action — run one whisper command — is unambiguous via the Quick start examples, so it qualifies for 5 under the simple-skill exception. Transcription is neither destructive nor a batch operation, so the validation cap does not apply.

5 / 5

Progressive Disclosure

The skill is under 50 lines with no need for external references, and the content is organized into clear Quick start and Notes sections, scoring 5 under the short-skill exception. No bundle files exist and the body makes no dangling file references, so there is nothing misfiled or nested.

5 / 5

Total

19

/

20

Passed

Description

53%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is concise and correctly third-person, with a clear and specific 'what', but it entirely lacks a 'when to use' clause and natural trigger vocabulary like 'transcribe' or 'audio'. It reads as a terse tool label rather than a trigger-rich description, which limits discoverability and completeness.

Suggestions

Add an explicit 'Use when...' clause, e.g., 'Use when the user wants to transcribe or translate audio files (mp3, m4a, wav) locally without an API key.'

Include natural trigger terms users would actually say — 'transcribe', 'transcription', 'audio', 'subtitles' — to improve trigger-term coverage.

Enumerate one or two more concrete capabilities (e.g., 'transcribe audio and translate to English, output txt/srt/vtt') to raise specificity.

DimensionReasoningScore

Specificity

"Local speech-to-text with the Whisper CLI" names the domain and one concrete capability but lists no multiple concrete actions (e.g., transcribe, translate, generate subtitles), matching the '1-2 concrete actions, not comprehensive' anchor. It is more concrete than the score-2 anchor but lacks the list of several specific actions required for 4.

3 / 5

Completeness

The description gives a clear 'what' (local speech-to-text via the Whisper CLI, no API key) but contains no 'Use when...' clause or equivalent trigger guidance, which caps completeness at 3 per the judging guidelines. It is well above anchor 2 because the 'what' is specific, not vague.

3 / 5

Trigger Term Quality

"speech-to-text" and "Whisper" are relevant keywords, but common natural phrases users would say — "transcribe", "transcription", "audio", "mp3" — are absent. Some relevant keywords with missing variations matches anchor 3, not the good synonym coverage of anchor 4.

3 / 5

Distinctiveness Conflict Risk

"Whisper CLI", "Local", and "(no API key)" carve out a mostly distinct niche with only minor overlap risk against other transcription/STT skills (e.g., cloud API-based ones). Not 5 because there are no explicit distinct trigger phrases, and generic 'speech-to-text' could still overlap similar skills.

4 / 5

Total

13

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

13

/

16

Passed

Repository
Bitterbot-AI/bitterbot-desktop
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.