CtrlK
BlogDocsLog inGet started
Tessl Logo

eval-harness

Formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles

51

Quality

56%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.agents/skills/eval-harness/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

63%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is well-sectioned with concrete templates and a clear four-phase workflow, but it repeats its own templates, leans on pseudocode for the actual evaluation step, and inlines everything monolithically instead of splitting examples and grader references into separate files. The /eval commands it documents don't exist in the bundle.

Suggestions

De-duplicate the eval-definition/report templates — present the format once and have the Workflow and Example sections reference it rather than re-printing it.

Make the Evaluate step executable: specify how a capability eval is actually run and recorded, and add an explicit failure loop (eval fails → fix → re-run → re-report).

Move the grader-type details and the full add-authentication example into reference files (e.g. references/graders.md, references/example-auth.md) and link them one level deep from SKILL.md.

DimensionReasoningScore

Conciseness

Mostly efficient templates, but the eval-definition and report formats are repeated three times (Eval Types, Workflow §Define, and the full 'Example: Adding Authentication'), and the Philosophy section restates EDD basics Claude can infer ('Define expected behavior BEFORE implementation', 'Run evals continuously').

3 / 5

Actionability

Concrete executable snippets are present (grep/npm-test grader commands, copy-pasteable eval templates), but the central Evaluate step is pseudocode ('[Run each capability eval, record PASS/FAIL]') and the '/eval define|check|report' commands have no implementation behind them.

4 / 5

Workflow Clarity

The Define → Implement → Evaluate → Report sequence is clear and evals themselves act as checkpoints, but there is no explicit failure-handling loop (what to do when an eval fails, fix, and re-run).

4 / 5

Progressive Disclosure

The body is ~230 lines with no bundle files at all; the grader-type guides, the integration patterns, and the full authentication example are inlined where they belong in separate reference files, though section headers keep it navigable.

3 / 5

Total

14

/

20

Passed

Description

48%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description names a clear niche but reads as a definition rather than a capability statement: no concrete actions, no natural trigger phrases, and no 'when to use' guidance. Users searching for eval or benchmarking help are unlikely to match it naturally.

Suggestions

List 2-3 concrete capabilities in the description, e.g. 'Define pass/fail criteria for agent tasks, run regression eval suites, and report pass@k reliability metrics.'

Add an explicit trigger clause: 'Use when setting up evals, benchmarking agent reliability, or checking regressions after prompt or agent changes.'

Include natural synonym keywords users would actually type — 'evals', 'benchmarking', 'pass@k', 'regression testing' — instead of only the 'eval-driven development (EDD)' jargon.

DimensionReasoningScore

Specificity

The description names the domain ('Formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles') but lists no concrete actions such as defining pass/fail criteria, running regression evals, or measuring pass@k — the domain is named while the actions remain generic.

2 / 5

Completeness

The 'what' is reasonably clear (a formal evaluation framework implementing EDD principles), but there is no 'Use when...' or equivalent trigger guidance, capping completeness at 3 per the judging guidelines.

3 / 5

Trigger Term Quality

Terms like 'evaluation framework', 'eval-driven development', 'EDD', and 'Claude Code sessions' are relevant but form a single jargon cluster; common natural variations users would say ('pass@k', 'benchmark', 'regression testing', 'grading') are missing.

3 / 5

Distinctiveness Conflict Risk

The EDD niche is mostly distinct ('eval-driven development', 'Claude Code sessions'), but 'evaluation framework' overlaps with general testing, benchmark, and CI-related skills, leaving minor conflict risk.

4 / 5

Total

12

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

Total

15

/

16

Passed

Repository
ysyecust/everything-claude-code
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.