CtrlK
BlogDocsLog inGet started
Tessl Logo

audio-transcription

Transcribe local audio/video and Apple Voice Memos quickly with cached MLX Whisper models, including bad/low-quality audio.

63

Quality

79%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/audio-transcription/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

82%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An efficient, highly actionable single-file skill: concrete commands, exact model IDs, hallucination-resistant flags, and a real validation checkpoint with red flags for rerun. The main weaknesses are that the fast path relies on helper scripts absent from the bundle and uses a user-specific absolute path, and the workflow sequence is implicit in the script rather than stated stepwise.

Suggestions

Ship the referenced './transcribe-audio.py' and './precache-models.py' in the bundle (e.g. a scripts/ directory) or demote them to optional helpers and make the manual uvx command the primary path, since the scripts are currently referenced but missing.

Replace the machine-specific absolute path '/Users/mitsuhiko/Development/agent-stuff/skills/audio-transcription' with a portable reference (e.g. 'from this skill's directory') so commands are copy-paste ready anywhere.

Make the run order explicit as a short numbered sequence (stage copy → verify cache → transcribe → check red flags → rerun/clean) with a fix-and-retry note per red flag, so the workflow does not depend on inferring steps from the script's bullet list.

DimensionReasoningScore

Conciseness

The body is lean and assumes competence: no explanation of what Whisper or transcription is, terse numbered rules ('Preserve temporary inputs immediately', 'Force language when known'), and every command block carries new information. The only mild redundancy (the manual template restating the rule-4 flags, and the repeated `cd` line) is justified as a self-contained fallback, so it fits 'every token earns its place' better than the 'minor instances of over-explanation' anchor at 4.

5 / 5

Actionability

Commands are concrete and copy-paste shaped: exact invocations ('uvx --from mlx-whisper mlx_whisper ... --hallucination-silence-threshold 2'), model IDs, output dirs, and a fully executable manual template. It falls short of 5 because the primary fast path depends on './transcribe-audio.py' and './precache-models.py', which are not present in the skill bundle, and the hardcoded absolute path '/Users/mitsuhiko/Development/agent-stuff/skills/audio-transcription' is machine-specific rather than portable.

4 / 5

Workflow Clarity

The sequence is clear (stage a stable copy → ensure the model is cached → run with hallucination-resistant flags → inspect output → rerun/clean up) and the 'Quality checks' section provides an explicit validation checkpoint with concrete red flags ('repeated phrases for many lines', 'avg_logprob is NaN'). It misses 5 because the ordering is partly implicit — the steps live inside the helper script's description rather than as an explicit numbered workflow with fix-and-retry instructions for each red flag.

4 / 5

Progressive Disclosure

A single, well-sectioned file (Core rules, Fast path, Cached models, Manual command template, Quality checks) at a size that justifies keeping everything inline, with the manual template correctly serving as fallback detail. Scored against the actual bundle structure, the minor gap is that the two referenced helper scripts ('./transcribe-audio.py', './precache-models.py') do not exist in the bundle, so the primary path's entry points are dangling references — matching 'good structure, minor organization gaps' rather than a 5.

4 / 5

Total

17

/

20

Passed

Description

65%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A tight, specific description with a clear niche and natural trigger vocabulary, but it omits any explicit 'when to use this' clause and under-represents the skill's actual outputs (formats, timestamps, hallucination cleanup). Adding a 'Use when...' sentence with common synonyms (dictation, meeting recording, lecture, mp3/m4a) would lift completeness and trigger coverage to the top anchors.

Suggestions

Add an explicit trigger clause, e.g. 'Use when the user asks to transcribe an audio or video file, a Voice Memos recording, dictation, a lecture, or a meeting, or mentions bad/hard-to-hear audio.'

Include common trigger synonyms and file extensions (dictation, meeting/lecture recording, mp3, m4a, wav) so the description matches the richer trigger list already present in the body.

Mention the concrete deliverables (txt/srt/vtt/json transcripts with timestamps and hallucination-resistant handling) so the 'what' is comprehensive, not just 'transcribe'.

DimensionReasoningScore

Specificity

The description names the domain ('Transcribe local audio/video and Apple Voice Memos') and one concrete action with a specific tool ('cached MLX Whisper models'), but stops at transcription — no mention of outputs (txt/srt/vtt/timestamps) or cleanup, so it matches '1-2 concrete actions, not comprehensive'. It is not a 4 because it does not list several distinct actions.

3 / 5

Completeness

The 'what' is clear (transcription with cached local MLX Whisper models, including bad audio), but there is no 'Use when...' clause or equivalent explicit trigger guidance; 'including bad/low-quality audio' describes scope, not when to invoke. Per the rubric guideline, a missing 'Use when' clause caps completeness at 3.

3 / 5

Trigger Term Quality

Natural user terms are present: 'transcribe', 'audio/video', 'Apple Voice Memos', 'bad/low-quality audio'. However common variations users would say — dictation, meeting/lecture recording, mp3/m4a — are absent, matching 'good keyword coverage; a few natural terms missing' rather than the comprehensive synonym-plus-extension coverage of a 5.

4 / 5

Distinctiveness Conflict Risk

The niche is unmistakable — local MLX Whisper transcription of audio/video and Apple Voice Memos — and 'MLX Whisper', 'Voice Memos', and 'bad/low-quality audio' are triggers no other plausible skill would claim. It clearly matches 'clear niche with distinct triggers; minimal conflict risk'.

5 / 5

Total

15

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
mitsuhiko/agent-stuff
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.