CtrlK
BlogDocsLog inGet started
Tessl Logo

eval-guide

Guide for running statistically meaningful agent-tty evals with trials, parallelism, and A/B comparison. Covers non-determinism baseline, recommended sample sizes, and result interpretation.

53

Quality

58%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

The risk profile of this skill

Fix and improve this skill with Tessl

tessl review fix ./.mux/skills/eval-guide/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

77%Scale 1-3

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

This is a strong, highly actionable eval guide with excellent executable examples and a well-structured A/B comparison workflow. Its main weakness is moderate verbosity in the framing/narrative sections and a monolithic structure that could benefit from splitting reference material (authoring, reporters, presets) into separate files. The statistical reasoning guidance and concrete threshold values are particularly valuable additions that Claude wouldn't know from general knowledge.

Suggestions

Tighten Section 1 to a bullet list of key numbers (flip rates, score change rates) without the narrative framing about 'what we learned' and 'the important point'.

Consider extracting Sections 7-10 (authoring, reporters, presets, snapshots) into separate reference files and linking to them from the main guide to improve progressive disclosure.

DimensionReasoningScore

Conciseness

The content is mostly efficient and information-dense, but includes some unnecessary framing (e.g., 'The short version' preamble, explaining what noise means, restating that single runs aren't trustworthy multiple times). Section 1's narrative about what 'we learned' could be tightened to just the key numbers and takeaway.

2 / 3

Actionability

Excellent actionability throughout: fully executable bash commands with realistic flags, specific trial count recommendations, concrete threshold values (0.05 score delta, 0.05 pass-rate delta), and copy-paste ready A/B comparison workflows. The quick reference commands section alone provides complete, runnable examples for every lane.

3 / 3

Workflow Clarity

The A/B comparison workflow in Section 3 is clearly sequenced with explicit steps (run baseline → save path → make change → run candidate with --compare-baseline → read verdicts). Section 4 provides clear interpretation rules with validation checkpoints (check if CI excludes 0, check effect size). The overall flow from understanding noise → running trials → comparing → interpreting is well-structured.

3 / 3

Progressive Disclosure

The content is well-organized with numbered sections and clear headers, but it's a long monolithic document (~200 lines) with no references to external files. Sections 7-10 cover distinct topics (authoring, reporters, presets, snapshots) that could be split into separate reference files. However, since no bundle files are provided, the inline approach is acceptable if not ideal.

2 / 3

Total

10

/

12

Passed

Description

40%Scale 1-3

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description identifies a clear and distinctive niche around agent-tty evaluation methodology, which is its main strength. However, it reads more like a topic summary than an actionable skill description, lacking concrete actions and entirely missing a 'Use when...' clause that would help Claude know when to select it. The trigger terms are somewhat technical and could miss natural user phrasings.

Suggestions

Add an explicit 'Use when...' clause, e.g., 'Use when the user wants to run agent-tty evals, compare agent performance, set up A/B tests, or interpret evaluation results.'

Replace 'Guide for running' and 'Covers' with specific concrete actions, e.g., 'Configures and runs statistically meaningful agent-tty evals, sets up A/B comparisons, determines sample sizes, and interprets trial results.'

Include natural trigger term variations users might say, such as 'benchmarks', 'testing agents', 'statistical significance', or 'compare runs'.

DimensionReasoningScore

Specificity

Names the domain (agent-tty evals) and mentions several concepts (trials, parallelism, A/B comparison, non-determinism baseline, sample sizes, result interpretation), but these read more like topic coverage than concrete actions. It says 'Guide for running' and 'Covers' rather than listing specific executable actions like 'configure trial parameters' or 'generate comparison reports'.

2 / 3

Completeness

The description addresses 'what' (guide for running evals with various features) but completely lacks a 'Use when...' clause or any explicit trigger guidance for when Claude should select this skill. Per the rubric, a missing 'Use when...' clause should cap completeness at 2, and since the 'what' is also somewhat vague ('Guide for' rather than concrete actions), this scores a 1.

1 / 3

Trigger Term Quality

Includes some relevant keywords like 'evals', 'A/B comparison', 'sample sizes', and 'agent-tty', but misses common natural variations users might say such as 'benchmarks', 'testing', 'evaluation framework', 'statistical significance', or 'compare models'. The term 'agent-tty' is quite specific/technical and may not match how users phrase requests.

2 / 3

Distinctiveness Conflict Risk

The description targets a very specific niche — 'agent-tty evals' with statistical methodology — which is unlikely to conflict with other skills. The combination of agent-tty, A/B comparison, and statistical eval concepts creates a distinct identity.

3 / 3

Total

8

/

12

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation11 / 11 Passed

Validation for skill structure

No warnings or errors.

Repository
coder/agent-tty
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.