CtrlK
BlogDocsLog inGet started
Tessl Logo

writing-livekit-scenarios

Creates and maintains the scenarios a LiveKit agent simulation runs, and wires the agent to consume them. Use when the user asks "what should I test", "generate simulation scenarios", "write scenarios for my agent", "add a scenario for X", "organize my scenario files", "my simulations are flaky", "the scenario hits my real database", "seed state per scenario", "it passed but booked the wrong thing", "grade the final state", or wants to stress-test a flow before shipping. Covers generating a baseline with the LiveKit scenario generator and refining it, writing the cases generation misses, phrasing simulated-user instructions so they steer reliably, splitting scenarios into sets across files, and the agent-side code that seeds deterministic state from a scenario and fails a run on its final state. To run them use running-livekit-simulations.

74

Quality

93%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

86%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a tightly written, well-sequenced guide that earns nearly top marks: it is token-efficient, deliberately delegates volatile details (schema, CLI flags, wiring code) to references and a docs-lookup skill, and its references are real, well-signaled, and exactly one level deep. The only soft spot is actionability/workflow verification: the concrete artifacts (a sample scenario, the agent-side wiring code, an explicit post-write validation step) live entirely in references or lookups, leaving small gaps between instruction and execution.

Suggestions

Inline one short example scenario (instructions + agent_expectations skeleton) in the 'Write instructions' section so the two instruction shapes have a concrete, copy-adaptable exemplar in-file.

Add an explicit validation checkpoint after authoring a set, e.g. 'run one scenario from the file and confirm it parses and grades before writing the rest', to close the workflow's verification loop in-file rather than only via references.

For 'Make the agent consume the scenario', inline a minimal detect-the-simulation code sketch (even 5 lines) alongside the four steps, since the wiring branch is the highest-risk code the skill directs the user to write.

DimensionReasoningScore

Conciseness

Lean and efficient: it explicitly refuses to restate volatile facts ('Field names and commands change, so this skill doesn't restate them' — delegating schema/CLI details to another skill), assumes Claude's competence throughout, and every paragraph delivers rule-plus-reason rather than background explanation. Even the short justificatory lines ('A file you haven't read isn't a test suite') earn their place by motivating a rule, so it fits the 'every token earns its place' anchor rather than the 4-anchor's trimmable over-explanation.

5 / 5

Actionability

Guidance is mostly concrete and executable: one runnable command ('lk agent simulate text -n 10'), a numbered baseline-refinement procedure, two defined instruction shapes, a numbered four-step agent-wiring procedure (Detect/Seed/Mock/Grade), and concrete rules ('Write absolute dates', 'Judge by outcome… Don't prescribe wording'). It falls short of the 5-anchor's copy-paste-ready coverage because the actual scenario-file format and agent-side wiring code are deferred to references and a docs lookup, leaving no in-file example of a scenario or its expected structure — but it is well above the 3-anchor's pseudocode/incomplete level since the deferral is deliberate and the steps are unambiguous.

4 / 5

Workflow Clarity

The end-to-end sequence is clear and coherently ordered (generate baseline → read agent code → ask what to probe → write instructions → split into sets → wire the agent → grow from failures → avoid bad tests), with numbered lists for the refinement and wiring procedures. Checkpoints exist (confirm flags with --help, 'read every scenario', 'how to confirm the wiring took effect', the generator-upload confirmation gate), so it exceeds the 3-anchor, but a few verifications are implicit or delegated to references rather than stated as explicit validate-then-proceed steps, keeping it below the 5-anchor's explicit feedback-loop pattern.

4 / 5

Progressive Disclosure

A clear overview with well-signaled, one-level-deep references: all three referenced files (references/risk-coverage.md, references/scenario-craft.md, references/connecting-the-agent.md) exist, each is cited in context where relevant ('Details and examples are in references/scenario-craft.md', 'covers it in full, including how to confirm the wiring took effect'), and a References section indexes them with one-line descriptions. This matches the anchor for clear overview with easy navigation, with no nested references or inlined content that belongs elsewhere.

5 / 5

Total

18

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is exemplary: third-person, concrete about what the skill does, exhaustive in natural trigger phrases (including failure-symptom phrasings like 'my simulations are flaky' and 'it passed but booked the wrong thing'), and explicitly scoped against sibling skills. Both what and when are answered with concrete language and no padding.

DimensionReasoningScore

Specificity

Lists multiple specific concrete actions with comprehensive coverage: 'Creates and maintains the scenarios a LiveKit agent simulation runs', 'generating a baseline with the LiveKit scenario generator and refining it', 'phrasing simulated-user instructions', 'splitting scenarios into sets across files', and 'the agent-side code that seeds deterministic state from a scenario and fails a run on its final state'. It is not a 4 because the action list spans the full task (create, refine, organize, wire to the agent) with no coverage gaps, matching the comprehensive anchor.

5 / 5

Completeness

Clearly and explicitly answers both what ('Creates and maintains the scenarios… and wires the agent to consume them', plus the enumerated 'Covers…' capabilities) and when (an explicit 'Use when the user asks…' clause with concrete trigger phrases). It also disambiguates execution by routing to 'running-livekit-simulations', so neither the 3-anchor (weak when) nor the 4-anchor (when could be more explicit) fits.

5 / 5

Trigger Term Quality

Comprehensive coverage of natural phrases a user would actually say, including indirect complaint forms: 'my simulations are flaky', 'the scenario hits my real database', 'it passed but booked the wrong thing', 'what should I test', 'add a scenario for X', 'seed state per scenario', 'grade the final state', 'stress-test a flow'. Below-anchor 4 ('a few natural terms missing') does not apply — synonyms and symptom-style triggers are all present.

5 / 5

Distinctiveness Conflict Risk

Clear niche (authoring/maintaining LiveKit simulation scenarios) with distinct triggers and minimal conflict risk; the closing 'To run them use running-livekit-simulations' explicitly fences off the closest sibling skill, and the authoring triggers ('write scenarios', 'add a scenario', 'organize my scenario files') are distinct from running/documentation ones. Not a 4, because the only overlap (running simulations) is already actively disambiguated rather than merely minor.

5 / 5

Total

20

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

Total

15

/

16

Passed

Repository
livekit/agent-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.