CtrlK
BlogDocsLog inGet started
Tessl Logo

deepgram-js-speech-to-text

Use when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram Speech-to-Text v1 (`/v1/listen`) for prerecorded or live audio transcription. Covers `client.listen.v1.media.transcribeUrl` / `transcribeFile` (REST) plus `client.listen.v1.createConnection()` / `connect()` (WebSocket). Use `deepgram-js-audio-intelligence` for summarize/sentiment/topics/diarize overlays, `deepgram-js-conversational-stt` for Flux turn-taking on `/v2/listen`, and `deepgram-js-voice-agent` for full-duplex assistants. Triggers include "transcribe", "speech to text", "STT", "listen.v1", "nova-3", "live transcription", and "websocket transcription".

72

Quality

87%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

82%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A strong, dense reference skill: all three transcription paths have complete executable examples, the API surface is precisely sourced, and the gotchas capture real, non-obvious SDK behavior. The main gaps are mild redundancy, absent error-handling guidance for the WebSocket flow, and an inline parameter catalog that would fit better in a separate reference file.

Suggestions

Remove the duplicated two-step socket-flow explanation (the paragraph after the WebSocket quick start repeats Gotcha 2) and condense or drop the 'Central product skills' install section to save tokens.

Add a short error-recovery snippet for the WebSocket path (handling 'Close'/'error' events and reconnecting after waitForOpen fails) so the workflow includes validation checkpoints.

Move the 'Key parameters / API surface' lists for REST and WSS into a reference file (e.g. reference.md alongside the existing layered pointers) and keep only the most common flags inline.

DimensionReasoningScore

Conciseness

The body is efficient with complete code examples and no basic-concept padding, but the two-step socket flow is stated twice ("The repo examples use the two-step socket flow: createConnection() → register handlers → connect() → waitForOpen()" and Gotcha 2 "Repo examples are two-stage for WSS"), and the closing "Central product skills" install section is promotional rather than instructional. These are minor trims, matching the 'efficient; minor instances of over-explanation' anchor rather than the lean 5.

4 / 5

Actionability

Fully executable, copy-paste-ready quick starts cover the three common cases (transcribeUrl, transcribeFile with multiple upload shapes, live WebSocket), backed by exact source paths ("src/api/resources/listen/resources/v1/client/Client.ts") and concrete commands like sendFinalize({ type: "Finalize" }). This matches the 'fully executable; specific examples cover the common cases' anchor.

5 / 5

Workflow Clarity

The WSS flow has a clear sequence (createConnection → register handlers → connect → waitForOpen → sendMedia → sendFinalize) and Gotchas 2–4 reinforce ordering and liveness requirements, but there is no error-recovery guidance (connection failure, socket close, retry). This sits at 'clear sequence with most checkpoints present; minor validation gaps' — not 3 (the sequence is explicit and the skill involves no destructive/batch operations) and not 5 (no feedback loops or checkpoints are given).

4 / 5

Progressive Disclosure

Good structure with a well-signaled layered API reference ("reference.md", canonical OpenAPI/AsyncAPI URLs, Context7, product docs) that is one level deep, and a clean example-file index. However, the 'Key parameters / API surface' section inlines a substantial parameter catalog that could live in a reference file, matching 'good structure; most content appropriately placed; minor organization gaps' rather than the cleanly split 5.

4 / 5

Total

17

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An exemplary description: it states explicit trigger conditions, names the exact SDK methods for both REST and WebSocket paths, and pre-emptively disambiguates all sibling Deepgram skills. The only weakness is slightly incomplete natural-synonym coverage in the trigger list.

DimensionReasoningScore

Specificity

It lists multiple specific concrete actions covering both modes — "client.listen.v1.media.transcribeUrl / transcribeFile (REST) plus client.listen.v1.createConnection() / connect() (WebSocket)" — with full coverage of the v1 surface, matching the 'multiple specific concrete actions; comprehensive coverage' anchor. It is well above the 4 anchor, which expects minor coverage gaps.

5 / 5

Completeness

It explicitly answers both questions: 'when' via "Use when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram Speech-to-Text v1 (/v1/listen) for prerecorded or live audio transcription" and 'what' via the named REST and WebSocket API calls, matching the 'clearly and explicitly answers both what AND when with concrete trigger phrases' anchor. Not 4, since the 'when' clause is fully explicit rather than improvable.

5 / 5

Trigger Term Quality

"Triggers include 'transcribe', 'speech to text', 'STT', 'listen.v1', 'nova-3', 'live transcription', and 'websocket transcription'" gives good natural coverage, but common variations like 'captions', 'ASR', 'audio to text', or 'realtime transcription' are missing — matching 'good keyword coverage; a few natural terms missing' rather than the comprehensive synonym coverage of the 5 anchor.

4 / 5

Distinctiveness Conflict Risk

It carves a clear niche (v1 /v1/listen STT in this JS SDK) and actively routes away the three nearest neighbors ("Use deepgram-js-audio-intelligence for summarize/sentiment/topics/diarize overlays, deepgram-js-conversational-stt for Flux turn-taking on /v2/listen, and deepgram-js-voice-agent for full-duplex assistants"), matching 'clear niche with distinct triggers; minimal conflict risk'.

5 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
deepgram/deepgram-js-sdk
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.