CtrlK
BlogDocsLog inGet started
Tessl Logo

deepgram-python-speech-to-text

Use when writing or reviewing Python code in this repo that calls Deepgram Speech-to-Text v1 (`/v1/listen`) for prerecorded or live audio transcription. Covers `client.listen.v1.media.transcribe_url` / `transcribe_file` (REST) and `client.listen.v1.connect` (WebSocket). Use this skill for basic ASR; use `deepgram-python-audio-intelligence` for summarize/sentiment/topics/diarize overlays, `deepgram-python-conversational-stt` for turn-taking v2/Flux, and `deepgram-python-voice-agent` for full-duplex assistants. Triggers include "transcribe", "live transcription", "speech to text", "STT", "listen endpoint", "nova-3", "listen.v1".

74

Quality

91%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

82%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A strong, highly actionable reference: executable examples for every usage pattern, explicit decision guidance, and Deepgram-specific gotchas Claude would not know. The main costs are a duplicated async section and an inlined parameter list that slightly bloat the token budget and blur the SKILL.md-as-overview role.

Suggestions

Merge the 'Async equivalents' section into 'Async / deferred result patterns §1' — both show await-based AsyncDeepgramClient usage, so one example suffices and saves ~10 lines of tokens.

Move the 'Key parameters' list and the interim/final flag semantics into the referenced reference.md (or a references/ file), keeping SKILL.md as the overview with a pointer.

Clarify the reference.md path (e.g., references/reference.md) so the layering is unambiguous when the skill bundle is installed standalone.

DimensionReasoningScore

Conciseness

The body is dense and Deepgram-specific (auth scheme caveat, interim/final flag semantics, gotchas), but the "Async equivalents" code block substantially duplicates "Async / deferred result patterns §1", which could be merged to save tokens. Efficient overall with one clear redundancy rather than pervasive padding.

4 / 5

Actionability

Every section ships copy-paste-ready executable code: auth setup, transcribe_url/transcribe_file (including the bytes-vs-iterator caveat), a complete threaded WSS loop with interim/final handling, async variants, and the callback/webhook pattern. The comparison table tells the reader exactly which pattern to pick.

5 / 5

Workflow Clarity

REST-vs-WSS-vs-callback selection is clearly guided ("When to use this product" section plus the pattern table), and the WSS teardown sequence is explicit ("send_finalize() to force final results... send_close_stream() after"), with an ERROR handler and gotchas covering failure modes. It stops short of explicit error-recovery/feedback loops, so it does not reach the top anchor.

4 / 5

Progressive Disclosure

The layered "API reference" section clearly signals one-level-deep resources (reference.md, canonical OpenAPI/AsyncAPI, Context7, docs), and the body stays a quick-start overview. However, reference.md is cited without a path and no bundle files exist alongside the skill, and the key-parameters list and flag-semantics detail could arguably live in that reference — good structure with minor gaps.

4 / 5

Total

17

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An exemplary description: concrete SDK methods and modes for 'what', an explicit 'Use when' clause with a natural trigger-term list for 'when', and explicit routing to sibling skills to avoid conflicts. No padding, no over-claims.

DimensionReasoningScore

Specificity

The description lists multiple concrete capabilities — "client.listen.v1.media.transcribe_url / transcribe_file (REST) and client.listen.v1.connect (WebSocket)" for "prerecorded or live audio transcription" — naming exact SDK methods and both transport modes with no vague filler.

5 / 5

Completeness

It explicitly answers 'when' ("Use when writing or reviewing Python code in this repo that calls Deepgram Speech-to-Text v1 (/v1/listen)") and 'what' (the specific methods and modes), plus a concrete trigger list — matching the top anchor exactly.

5 / 5

Trigger Term Quality

"Triggers include 'transcribe', 'live transcription', 'speech to text', 'STT', 'listen endpoint', 'nova-3', 'listen.v1'" covers natural user phrasings, abbreviations, and technical synonyms, reinforced by 'basic ASR' and 'prerecorded or live audio transcription' elsewhere in the text.

5 / 5

Distinctiveness Conflict Risk

It carves a clear niche (basic ASR via /v1/listen) and explicitly disambiguates sibling skills — "use deepgram-python-audio-intelligence for summarize/sentiment/topics/diarize overlays, deepgram-python-conversational-stt for turn-taking v2/Flux, and deepgram-python-voice-agent for full-duplex assistants" — minimizing wrong-skill triggering.

5 / 5

Total

20

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
deepgram/deepgram-python-sdk
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.