Writes turn-level tests for a LiveKit agent in the user's normal test suite: pytest (Python) or Vitest (Node.js). Use when the user asks to "write tests for my agent", "add a test for this tool", "test the handoff", "pin this bug", "why does my agent test fail", or after building or changing agent behavior that needs regression coverage. Covers the SDK's test session harness, assertions on messages, tool calls and handoffs, LLM judging of a reply against an intent, mocking tools, multi-turn tests, and judging whole conversations with the built-in judges. For interactive poking use debugging-livekit-agents. For grading whole conversations at scale use running-livekit-simulations.
73
92%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
Passed
No findings from the security scan
Turn-level tests are the cheapest lasting verification an agent can have. They run in the user's
existing suite, in text mode, and they're fast enough for every commit. The framework's helper
names and signatures change, so look them up with reading-livekit-docs before writing, and use
the testing docs page for every API detail this skill leaves out.
Every test has the same shape: start a test session with the agent under test, run one user turn, and assert on the events that turn produced.
A turn produces a sequence of events. A simple turn is one message. A more typical one is a tool call, its output, maybe a handoff, and then a message. You write the test by walking that sequence in order, asserting on each event, and then asserting the turn has nothing more.
There are three kinds of assertion, and most of this skill is knowing which one to use.
Structural assertions check message roles, that a tool was called, its arguments, what it returned, and that a handoff to a specific agent happened. They're deterministic and fail for exactly one reason, so prefer them.
Judged assertions use an LLM-judge helper that gives one message and an intent string to a model and asks whether they match. Use them for the content of a reply, which you can't assert exactly. Describe the intent by outcome ("tells the user the booking is confirmed and gives the time"), not by wording.
Whole-conversation assertions use a judge-group helper that runs several built-in judges concurrently over the whole chat history and aggregates their verdicts. The built-in judges cover dimensions like grounding, relevance, safety, task completion, and tool use; the docs list the current set. Use this when the question spans several turns.
Rules of thumb:
Tests shouldn't hit real backends, since that makes a suite slow and nondeterministic. Override the tools for the agent under test and return fixed values.
Two things beyond the API:
run() calls, which is what
a test needs. A session-scoped form also exists for a session that runs on its own and needs
mocks active for its whole lifetime. Simulation entrypoints use that form, tests don't. See
writing-livekit-scenarios.Mocking only changes execution. The model still sees the real tool schemas, so tool selection is still under test.
There are two ways to test behavior that depends on earlier turns:
Roughly in order of value:
A test suite for an agent that does things — books, edits, confirms — needs three kinds of test, and confusing them is how a green suite ships a broken agent:
Shortcuts that look like tests and aren't: calling a helper or filling private state instead of driving the interaction; telling the test caller which tool to name; mocking the very mutation whose correctness is under test; comparing output only with the application's own exporter; retrying a failed turn until it passes; deleting the assertion that failed.
The most common bad agent test asserts that the agent does something it shouldn't: states data it can't know, gives a specific medical, legal, or financial recommendation, or completes a flow that should have been blocked. If the agent can only pass by misbehaving, fix the test.
For a guardrail, the pass is that the agent refuses, escalates, or declines to make something up. Write the assertion that way.
debugging-livekit-agents. Faster loop, no assertions kept.running-livekit-simulations.
More expensive, and catches emergent behavior that turn-level tests miss.When a simulation keeps failing the same way, the bug is usually turn-level. Write a test for it here, which pins down the cause more precisely and catches it earlier.
reading-livekit-docsdebugging-livekit-agentswriting-livekit-scenarios, running-livekit-simulations5d7488b
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.