CtrlK
BlogDocsLog inGet started
Tessl Logo

eval-harness

Eval-driven development (EDD) framework for AI coding sessions — define capability and regression evals before coding, grade with code-based, model-based, rule, or human graders, and track pass@k and pass^k reliability. Use when defining pass/fail criteria for agent tasks, measuring agent reliability, building regression suites for prompt or agent changes, or benchmarking across model versions.

64

Quality

78%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/eval-harness/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

63%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a well-templated, largely actionable EDD guide, but it is held back by duplication (grader types, pass@k guidance, and storage layout each appear twice), placeholder steps in the evaluate workflow, and progressive-disclosure problems: the ~300-line monolith inlines detail that belongs in referenced files, and its file references point to bundle paths that are absent. Consolidating the duplicates and either shipping or removing the dangling references would lift it substantially.

Suggestions

Deduplicate the repeated material: merge 'Product Evals (v1.8)' grader types and pass@k guidance into the earlier 'Grader Types' and 'Metrics' sections, and collapse 'Eval Storage' with 'Minimal Eval Artifact Layout'.

Add an explicit failure-handling loop to the Evaluate phase (e.g. 'if any eval fails: fix the regression, re-run `/eval check <feature>`, only report when all pass'), and replace placeholder steps like '[Run each capability eval, record PASS/FAIL]' with concrete commands.

Fix the reference topology: either ship the referenced files (scripts/eval-harness.js, scripts/lib/eval-harness/, docs/architecture/eval-harness-frameworks.md) or remove the pointers, and move the dense 'Local Framework Utilities' prose into a one-level-deep reference file to slim SKILL.md.

DimensionReasoningScore

Conciseness

Mostly efficient — templates, commands, and formats rather than concept explanations — but several sections are padded or duplicated: 'When to Activate' restates the frontmatter triggers, the grader-type list appears twice ('Grader Types' and again under 'Product Evals (v1.8)' as 'Code grader / Rule grader / Model grader / Human grader'), pass@k guidance appears twice ('Metrics' and 'pass@k Guidance'), and the eval storage layout is described twice ('Eval Storage' and 'Minimal Eval Artifact Layout'). This fits the level-3 anchor ('mostly efficient but includes some unnecessary explanation or could be tightened'); it is not level 4 because the duplication is substantive rather than minor, but not level 2 since nothing explains concepts Claude wouldn't know.

3 / 5

Actionability

The body provides mostly executable guidance: copy-paste grader commands ('grep -q "export function handleAuth" src/auth.ts && echo "PASS" || echo "FAIL"'), CLI invocations ('node scripts/eval-harness.js example', '/eval check feature-name'), and complete markdown eval templates. It is not level 5 because some steps remain placeholders, e.g. '[Run each capability eval, record PASS/FAIL]' and '[Write code]' in the workflow example, leaving minor gaps for a fresh user.

4 / 5

Workflow Clarity

The four-phase sequence (Define → Implement → Evaluate → Report) is clearly laid out with a worked 'Example: Adding Authentication', and regression evals act as a built-in checkpoint. It is not level 5 because the Evaluate phase lacks an explicit failure feedback loop — there is no 'if an eval fails, fix and re-run' step, and the workflow example defers to '/eval check' rather than spelling out validation; it is well above level 3 since sequence and most checkpoints are present.

4 / 5

Progressive Disclosure

The body is sectioned with clear headers, but the bundle contains no references/, scripts/, or assets/ directories, so the inline pointers ('scripts/lib/eval-harness/', 'node scripts/eval-harness.js example', 'See docs/architecture/eval-harness-frameworks.md') reference files that do not exist in the skill, and substantial material (framework utility details, the duplicated Product Evals v1.8 content) is inlined where it could live one level deep. This matches the level-3 anchor ('some structure but could be better organized; references present but not clearly signaled; content that should be separate is inline') — structure exists, but the reference topology is broken and content placement needs work.

3 / 5

Total

14

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: it states concrete capabilities in third person, includes an explicit 'Use when' clause with natural trigger phrases, and carves out a distinctive niche. The only improvement would be broadening trigger-term coverage with a few more synonyms users might naturally say.

Suggestions

Add a few natural synonyms to the trigger clause (e.g. 'writing evals for agents', 'testing whether an agent change regressed behavior', 'flaky graders') to reach comprehensive trigger coverage.

Clarify that the workflow applies to any AI coding agent session, not only Claude Code, if that is the intended scope.

DimensionReasoningScore

Specificity

The description lists multiple concrete actions — 'define capability and regression evals before coding', 'grade with code-based, model-based, rule, or human graders', 'track pass@k and pass^k reliability' — giving comprehensive coverage of the skill's capabilities. It matches the level-5 anchor ('lists multiple specific concrete actions; comprehensive coverage') and is not the level-4 anchor because there are no meaningful gaps in the action inventory.

5 / 5

Completeness

It explicitly answers both questions: 'what' (define capability and regression evals, grade with four grader types, track pass@k and pass^k) and 'when' via an explicit 'Use when' clause with concrete trigger phrases. This directly matches the level-5 anchor example's structure and is not level 4, where the 'when' would be less specific.

5 / 5

Trigger Term Quality

Strong natural phrases users would say: 'defining pass/fail criteria', 'measuring agent reliability', 'building regression suites', 'benchmarking across model versions'. It falls short of the level-5 anchor because common synonyms and variations are missing (e.g. 'testing an agent change', 'evals', 'flaky tests'), so it sits between the good-coverage and comprehensive-coverage anchors, noticeably above the midpoint.

4 / 5

Distinctiveness Conflict Risk

The niche is clear — eval-driven development for AI coding sessions with pass@k/pass^k metrics and grader taxonomies — and the trigger terms ('pass/fail criteria for agent tasks', 'benchmarking across model versions') are unlikely to fire for unrelated skills. It matches the level-5 'clear niche with distinct triggers' anchor; the only conceivable overlap is generic test-suite work, which the EDD/pass@k framing disambiguates.

5 / 5

Total

19

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

referenced_paths_exist

Referenced path issues: 3 missing, 1 deeper-than-1-level

Warning

Total

13

/

16

Passed

Repository
affaan-m/ECC
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.