CtrlK
BlogDocsLog inGet started
Tessl Logo

eval-guide

Guide for running statistically meaningful agent-tty evals with trials, parallelism, and A/B comparison. Covers non-determinism baseline, recommended sample sizes, and result interpretation.

64

Quality

76%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.mux/skills/eval-guide/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

82%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The content is highly actionable and well-structured, with executable commands and concrete statistical guidance throughout. It is concise for its density, though a few narrative sections and the inline command catalog leave minor room for tighter organization.

Suggestions

Consider moving the repeated quick-reference command blocks (section 6) into a separate references file to keep SKILL.md as an overview.

Tighten the narrative in section 1 to bullet-only findings, dropping restated context Claude can infer.

Make the before/after workflow's validation gate explicit (e.g. 'Only declare improved when the paired CI excludes 0 and the delta >= 0.05') to strengthen the checkpoint.

DimensionReasoningScore

Conciseness

The body is information-dense with concrete numbers and assumes Claude's competence, avoiding explanations of known concepts; a few narrative passages (e.g. the findings recap) could be trimmed slightly.

4 / 5

Actionability

Provides copy-paste ready, fully executable bash commands with concrete flags, specific trial/concurrency counts, and precise practical cutoffs (0.05 deltas) covering the common cases.

5 / 5

Workflow Clarity

The before/after comparison is a clearly numbered sequence with validation rules (CI excludes 0, effect-size cutoff) and a feedback loop (add trials when noise-dominated); minor checkpoints are implicit rather than gated.

4 / 5

Progressive Disclosure

Content is well-organized into ten numbered sections with clear headers and no nested references; since no bundle files exist everything is inline, and some material (quick-reference commands, findings data) could be split out but is reasonable inline.

4 / 5

Total

17

/

20

Passed

Description

70%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description clearly communicates the skill's purpose and niche, with good specificity and trigger terms. Its main weakness is the missing explicit 'Use when...' trigger clause, which caps completeness.

Suggestions

Add an explicit 'Use when...' clause, e.g. 'Use when deciding whether a prompt or skill change actually improved agent-tty eval results.'

Include a few synonyms/variations users might say (e.g. 'eval runs', 'A/B test', 'paired comparison') to broaden trigger coverage.

Surface a couple more discrete actions (e.g. 'run paired baseline comparisons', 'set trial counts and concurrency') to push specificity toward comprehensive.

DimensionReasoningScore

Specificity

Names the domain and several concrete capabilities (trials, parallelism, A/B comparison, recommended sample sizes, result interpretation), with only minor coverage gaps.

4 / 5

Completeness

The 'what' is clear, but there is no explicit 'Use when...' trigger guidance, so completeness is capped at 3 per the rubric guideline.

3 / 5

Trigger Term Quality

Includes natural terms a user would say (agent-tty evals, trials, A/B comparison, sample sizes), though it lacks a 'Use when' trigger phrase and broader synonyms.

4 / 5

Distinctiveness Conflict Risk

Tied to a clear niche ('agent-tty evals') with distinct triggers and minimal overlap risk with other skills.

5 / 5

Total

16

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
coder/agent-tty
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.