CtrlK
BlogDocsLog inGet started
Tessl Logo

ax-audio

This skill helps an LLM generate correct audio code with @ax-llm/ax. Use when the user asks about ai.transcribe(), ai.speak(), signature audio inputs or outputs, agent audio behavior, .chat() conversational audio, OpenAI audio or realtime models, Gemini Live native audio, Grok Voice Agent models, voices, formats, transcripts, or how audio fits with structured outputs.

68

Quality

85%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

78%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

This is a highly actionable, information-dense codegen reference with executable examples for every audio surface and accurate decision routing up front. Its main weakness is structural: everything lives inline in SKILL.md with zero progressive disclosure, so provider-specific deep-dives (especially Meta Muse and the default-config listings) should be split into one-level-deep reference files.

Suggestions

Move the per-provider deep-dives (Meta Muse Voice, Gemini Live, Grok Voice, and the OpenAI default-config listings) into one-level-deep files under references/ (e.g., references/meta-muse.md), keeping SKILL.md as an overview with clearly signaled links per surface.

Tighten the Meta Muse prose into short bullet rules (sample-rate agreement, isFinal snapshot replacement, endStream close semantics) to trim the most padded section.

Add a short 'Choosing a surface' decision list with error-recovery notes (e.g., what to do when a provider throws AxMediaNotSupportedError) to consolidate the sequencing guidance currently spread across sections.

DimensionReasoningScore

Conciseness

The body is dense with library-specific facts Claude cannot know (default models, voices, mime-type behavior, WebSocket requirements) and avoids explaining generic concepts, so most tokens earn their place. However, sections like the Meta Muse prose ("Ax folds cumulative or delta partials, overlapping turns, speaker labels, progress events, errors, and graceful endStream completion...") and the full inline Config Shape type could be tightened, fitting anchor 4 (efficient with minor trims) rather than anchor 5. It is clearly above anchor 3, which would require noticeably unnecessary explanation.

4 / 5

Actionability

Every section provides complete, executable TypeScript with imports, constructor setup, model/voice parameters, and even streaming-result guards ("if (!('results' in res)) throw new Error('Expected a non-streaming chat response')"). The examples cover the common cases — batch transcribe/speak, signature artifacts, agent audio, chat audio, realtime WebSocket setup, and streaming deltas — matching anchor 5's copy-paste-ready coverage. Anchor 4 would require minor gaps in the code, which are not present.

5 / 5

Workflow Clarity

The body opens with an unambiguous decision guide ("Pick the smallest audio surface that matches the job") routing to transcribe/speak/signature/chat, and each surface has its own section with concrete setup steps (e.g., passing a WebSocket constructor for realtime). No validation checkpoints are needed since nothing is destructive or batch-mutating, but the multi-part setups (realtime history references, Meta stream lifecycle) are spread across sections without explicit sequencing or error-recovery loops, keeping it at anchor 4 rather than 5.

4 / 5

Progressive Disclosure

There are no bundle files at all (no references/, scripts/, or assets/), and the entire ~460-line provider reference — OpenAI defaults, Gemini Live, Grok Voice, and the lengthy Meta Muse guide — is inlined in SKILL.md, which is content that clearly belongs in separate per-provider reference files. Section headers are well organized and navigable, matching anchor 3's 'some structure but content that should be separate is inline', while anchor 2's example (300+ lines with no section headers) undervalues the good in-file organization.

3 / 5

Total

16

/

20

Passed

Description

87%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: it states what the skill does and gives an explicit, keyword-rich 'Use when' clause covering batch, signature, agent, and conversational audio across named providers. The only weaknesses are a missing Meta/Muse surface in the capability list and absent natural synonyms like 'text-to-speech'/'TTS'.

DimensionReasoningScore

Specificity

The description names concrete surfaces — "ai.transcribe(), ai.speak(), signature audio inputs or outputs, agent audio behavior, .chat() conversational audio" — which is specific and nearly comprehensive, but it omits the Meta/Muse Voice surface and streaming that the body covers, leaving minor gaps in coverage. It fits anchor 4 (several specific actions, minor gaps) better than anchor 5 (comprehensive coverage), and far better than anchor 3, which would require only 1-2 concrete actions.

4 / 5

Completeness

It explicitly answers 'what' ("helps an LLM generate correct audio code with @ax-llm/ax") and 'when' with a concrete trigger clause ("Use when the user asks about ai.transcribe(), ai.speak(), ... voices, formats, transcripts, or how audio fits with structured outputs"). This matches anchor 5 exactly; anchor 4 would require the 'when' to be less explicit, which it is not.

5 / 5

Trigger Term Quality

Trigger terms are natural and varied: "voices, formats, transcripts", "OpenAI audio or realtime models", "Gemini Live native audio", "Grok Voice Agent models", "structured outputs". A few natural synonyms users would say are missing — notably "text-to-speech", "speech-to-text", "TTS", and "realtime/voice assistant" phrasing — so it matches anchor 4 rather than anchor 5's comprehensive synonym coverage. It is well above anchor 3, since keyword coverage is broad, not just 'some relevant keywords'.

4 / 5

Distinctiveness Conflict Risk

The niche is tightly scoped to audio codegen with @ax-llm/ax, and triggers name provider-specific audio models and Ax APIs, so it is unlikely to fire for unrelated skills. This matches anchor 5 (clear niche with distinct triggers); anchor 4 would imply meaningful overlap with closely related skills, which the specific API-name triggers preclude.

5 / 5

Total

18

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
ax-llm/ax
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.