CtrlK
BlogDocsLog inGet started
Tessl Logo

eval-harness

Formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles. Use when a Claude Code workflow needs a formal eval before it is trusted or changed.

55

Quality

63%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.agents/skills/eval-harness/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

63%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body delivers a clear four-phase EDD workflow with concrete templates and executable grader examples. Its weaknesses are moderate duplication (repeated report formats and a worked example that restates earlier content), placeholder pseudo-steps in the Evaluate section, and a monolithic single-file structure with no progressive disclosure.

Suggestions

Remove the duplicated EVAL REPORT template — show it once and have the 'Adding Authentication' example reference it rather than reprint it.

Replace placeholder lines like '[Run each capability eval, record PASS/FAIL]' with actual executable commands, and either define the /eval define|check|report commands as real scripts or explain how they are dispatched.

Split the grader-type catalogue and the full worked example into reference files (e.g., references/graders.md, references/example-auth.md) linked one level deep from SKILL.md, and add an explicit fail → fix → re-evaluate feedback loop to the workflow.

DimensionReasoningScore

Conciseness

Mostly efficient with compact templates, but there is real duplication: the EVAL REPORT format appears twice (the 'feature-xyz' report and again in the 'Adding Authentication' example) and the 'Philosophy' section explains EDD conceptually. Anchor 3 ('mostly efficient but includes some unnecessary explanation or could be tightened') fits better than anchor 4 given this goes beyond minor trimming.

3 / 5

Actionability

Concrete eval-definition templates, real executable graders (grep/npm test commands), slash-command integration, and a storage layout give mostly executable guidance. Kept below 5 by placeholder lines like '[Run each capability eval, record PASS/FAIL]' and the '/eval define|check|report' commands that are referenced but never implemented or defined as scripts.

4 / 5

Workflow Clarity

The Define → Implement → Evaluate → Report sequence is clearly presented with PASS/FAIL recording and regression checks as checkpoints. Not 5 because the error-recovery feedback loop (eval fails → fix → re-run) is implicit rather than an explicit step.

4 / 5

Progressive Disclosure

The skill is a single 237-line file with well-organized sections but no external references at all, and content that could be split (the grader-type catalogue, the full worked authentication example) is inlined. This exceeds the under-50-line simple-skill exception, so it falls to anchor 3 ('some structure but could be better organized').

3 / 5

Total

14

/

20

Passed

Description

62%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A reasonably strong description with an explicit 'Use when' trigger and a distinct niche, written in third person. Its main weakness is the absence of concrete, enumerated capabilities — the 'what' is stated abstractly rather than through specific actions like defining eval criteria or measuring pass@k.

Suggestions

Enumerate 2-3 concrete capabilities in the description (e.g., 'Define pass/fail criteria, run capability and regression evals, and measure agent reliability with pass@k metrics').

Add natural trigger variations users might say, such as 'testing agent changes', 'regression suite', or 'benchmarking models'.

Make the 'what' more concrete than 'formal evaluation framework' by naming the artifacts it produces (eval definitions, run logs, eval reports).

DimensionReasoningScore

Specificity

The description names the domain ("Formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles") but offers only generic actions — no concrete verbs like defining pass/fail criteria or measuring pass@k. It sits above anchor 2 ('actions minimal or generic') because the EDD framing adds substance, but below anchor 4 because no specific action list is present.

3 / 5

Completeness

Both parts are answered: the 'what' ("Formal evaluation framework... implementing eval-driven development (EDD) principles") and an explicit 'when' ("Use when a Claude Code workflow needs a formal eval before it is trusted or changed"). Not 5 because the 'what' is abstract and the trigger phrasing is less concrete than rubric exemplars; not 3 because the 'when' clause is explicit, not merely implied.

4 / 5

Trigger Term Quality

Natural keywords present ("eval", "evaluation", "EDD", "needs a formal eval before it is trusted or changed") but missing common variations users would say such as "testing", "regression", or "benchmark". Matches anchor 3 ('some relevant keywords but missing common variations or synonyms').

3 / 5

Distinctiveness Conflict Risk

The EDD/formal-eval niche for Claude Code sessions is clearly distinguishable from most skills. Only minor overlap risk with general testing or code-review skills, matching anchor 4 ('mostly distinct; minor overlap risk with closely related skills').

4 / 5

Total

14

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

Total

15

/

16

Passed

Repository
affaan-m/ECC
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.