CtrlK
BlogDocsLog inGet started
Tessl Logo

testing-livekit-agents

Writes turn-level tests for a LiveKit agent in the user's normal test suite: pytest (Python) or Vitest (Node.js). Use when the user asks to "write tests for my agent", "add a test for this tool", "test the handoff", "pin this bug", "why does my agent test fail", or after building or changing agent behavior that needs regression coverage. Covers the SDK's test session harness, assertions on messages, tool calls and handoffs, LLM judging of a reply against an intent, mocking tools, multi-turn tests, and judging whole conversations with the built-in judges. For interactive poking use debugging-livekit-agents. For grading whole conversations at scale use running-livekit-simulations.

73

Quality

92%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

82%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A tightly written, opinionated body that assumes Claude's competence and consistently teaches judgment rather than basics. The gaps are mechanical rather than conceptual: no executable example to anchor the test shape, and all guidance inlined in one file rather than split across references.

Suggestions

Include one minimal executable example (e.g., a short pytest snippet of a session-start, one turn, and a structural assertion) marked as 'check current signatures via reading-livekit-docs' — it would anchor the described shape in concrete code without going stale.

Consolidate the placement framing: the intro and the 'Where this sits' section both argue that these tests are cheap, fast, and run every commit; stating it once would trim a few lines.

Consider moving the 'Three kinds of evidence, kept distinct' deep-dive into a references/ file, keeping SKILL.md as a leaner overview that points to it — this would complete the progressive-disclosure structure the other sections already gesture toward.

DimensionReasoningScore

Conciseness

The body never explains concepts Claude already knows (no 'what LiveKit is', no pytest tutorial) and every section teaches non-obvious judgment: 'Mock failures as well as successes', 'Mocking only changes execution. The model still sees the real tool schemas', the three-kinds-of-evidence taxonomy. The only trimmable material is slight framing duplication between the intro and 'Where this sits', which stays on the 5 anchor ('every token earns its place') rather than dropping to 4.

5 / 5

Actionability

The guidance is concrete and specific — the test shape ('start a test session with the agent under test, run one user turn, and assert on the events that turn produced'), a worked intent example ('tells the user the booking is confirmed and gives the time'), and a prioritized what-to-test list — but it contains no executable code or commands at all; all API mechanics are deferred to 'reading-livekit-docs' and 'the testing docs page'. That deferral is justified (helper names change), and the rubric's instruction-only note says not to penalize absent code when guidance is actionable, so it sits above the 3 anchor (guidance is complete, not pseudocode) but below 5 (nothing is copy-paste ready).

4 / 5

Workflow Clarity

A clear sequence is given: look up the API surface first ('look them up with reading-livekit-docs before writing'), start a session, run one turn, 'walking that sequence in order, asserting on each event', then 'Close the turn.' The anti-pattern sections provide error guidance ('If the agent can only pass by misbehaving, fix the test'), but there is no explicit validate/fix/retry feedback loop for a failing test run, which keeps it below the 5 anchor while clearly above 3.

4 / 5

Progressive Disclosure

The file is well-sectioned with clear headers, an explicit pointer for detail ('use the testing docs page for every API detail this skill leaves out'), and a closing 'Related skills' section — all references are clearly signaled and one level deep. However, all ~125 lines of guidance live inline in a single file with no bundled reference split (no references/ or scripts/ exist), and the under-50-lines simple-skill exception does not apply, so it falls just short of the 5 anchor's 'content appropriately split'.

4 / 5

Total

17

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An exemplary description: concrete capabilities, quoted natural-language trigger phrases, explicit what/when, and active disambiguation against sibling skills. All dimensions sit firmly at the top anchor.

DimensionReasoningScore

Specificity

Names the domain ('Writes turn-level tests for a LiveKit agent') and lists multiple concrete capabilities — 'test session harness, assertions on messages, tool calls and handoffs, LLM judging of a reply against an intent, mocking tools, multi-turn tests, and judging whole conversations with the built-in judges' — plus concrete tooling (pytest/Vitest). Coverage is comprehensive with no gaps, so it matches the 5 anchor rather than 4.

5 / 5

Completeness

Explicitly answers both: what — 'Writes turn-level tests for a LiveKit agent in the user's normal test suite: pytest (Python) or Vitest (Node.js)' — and when — a full set of concrete user-facing trigger phrases plus a post-change trigger. It matches the 5 anchor's structure exactly; the 'Use when' clause is present and specific, so the 3-cap does not apply.

5 / 5

Trigger Term Quality

Includes natural phrases users would actually say, quoted directly: '"write tests for my agent"', '"add a test for this tool"', '"test the handoff"', '"pin this bug"', '"why does my agent test fail"', plus the proactive trigger 'after building or changing agent behavior that needs regression coverage'. This is comprehensive natural-term coverage, a clear 5 rather than 4.

5 / 5

Distinctiveness Conflict Risk

Clear niche (turn-level tests in the user's own suite) with explicit routing away from adjacent skills: 'For interactive poking use debugging-livekit-agents. For grading whole conversations at scale use running-livekit-simulations.' This active disambiguation gives minimal conflict risk.

5 / 5

Total

20

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

Total

15

/

16

Passed

Repository
livekit/agent-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.