CtrlK
BlogDocsLog inGet started
Tessl Logo

running-livekit-simulations

Runs LiveKit agent simulations and acts on the results. Use when the user says "run my simulations", "regression test my agent before deploying", "run the scenarios", "use lk agent simulate", "did my agent pass", "why did this scenario fail", "run simulations in CI", "test the audio pipeline", "check turn-taking and interruptions", or wants to check whole-conversation behavior before shipping. Covers text and audio mode and what each catches, running against a local or deployed agent, degraded-audio flags, automating a pre-release run, and triaging failures with list, view and export. For writing the scenarios use writing-livekit-scenarios. Not the default for a bare "test my agent", which goes to debugging-livekit-agents. Use this skill when the user names simulations, scenarios, a run, CI, or shipping.

75

Quality

94%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A tight, high-signal body: it assumes Claude's competence, gives a runnable command with a deliberate deferral of volatile flags to --help, and provides an explicit triage-and-re-run feedback loop. The only soft spot is the occasional high-level hint ('an option lets you grade an already-running agent') where the specific flag would make guidance fully copy-paste ready.

Suggestions

Name the option/flag for grading an already-running agent (e.g. 'lk agent simulate --agent <name> text --scenarios scenarios.yaml') so that path is copy-paste ready rather than a hint.

For the export/list workflow, include one concrete example invocation (e.g. 'lk agent simulate export <run-id>') since '--help names the flags' currently leaves the reader to discover the exact syntax mid-task.

DimensionReasoningScore

Conciseness

Lean and efficient throughout: no explanation of concepts Claude already knows, no padding, and every sentence carries skill-specific information ('Text is the default... faster, cheaper and more deterministic'; 'The gap between the two is what the user experiences'). Deferring flags to '--help' avoids restating volatile details.

5 / 5

Actionability

Mostly executable guidance: a copy-paste command ('lk agent simulate text --scenarios scenarios.yaml'), named subcommands (list, export), and a concrete triage decision list. However, 'An option lets you grade an already-running agent by name' names neither the option nor its flag, leaving a hint where a specific invocation belongs.

4 / 5

Workflow Clarity

Clear sequence with explicit validation and feedback loops: triage failures into three categorized causes with fixes, then 'After a fix, run the whole file, not only the scenario you were working on', plus 'exits non-zero when any scenario fails' for CI and 'Move repeat failures down the stack' for escalation. Recovery paths are explicit and checklist-like.

5 / 5

Progressive Disclosure

Well-organized single-file skill with clear section headers and well-signaled pointers to sibling skills (reading-livekit-docs for flags, writing-livekit-scenarios for authoring). It references no detail files of its own, and at ~103 lines it exceeds the under-50-line simple-skill case, so structure is good but there is no one-level-deep reference organization to reward.

4 / 5

Total

18

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An exemplary description: third-person voice, comprehensive natural trigger phrases, explicit what-and-when, and proactive disambiguation from sibling skills in the same family. Every clause earns its place with no fluff or over-claiming.

DimensionReasoningScore

Specificity

The description lists multiple specific concrete actions with comprehensive coverage: 'Runs LiveKit agent simulations and acts on the results', 'triaging failures with list, view and export', 'degraded-audio flags, automating a pre-release run', covering both text and audio modes. No gaps in capability coverage are evident.

5 / 5

Completeness

Explicitly answers both questions: what ('Runs LiveKit agent simulations and acts on the results... Covers text and audio mode... triaging failures with list, view and export') and when ('Use when the user says...'). Both are concrete and fully explicit with trigger phrases.

5 / 5

Trigger Term Quality

Comprehensive natural trigger phrases users would actually say: 'run my simulations', 'regression test my agent before deploying', 'did my agent pass', 'why did this scenario fail', 'run simulations in CI', 'check turn-taking and interruptions', plus the exact CLI form 'use lk agent simulate'.

5 / 5

Distinctiveness Conflict Risk

Clear niche with explicit disambiguation from sibling skills: 'For writing the scenarios use writing-livekit-scenarios. Not the default for a bare "test my agent", which goes to debugging-livekit-agents.' Minimal conflict risk.

5 / 5

Total

20

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

Total

15

/

16

Passed

Repository
livekit/agent-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.