CtrlK
BlogDocsLog inGet started
Tessl Logo

eval-harness

Formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles

82

2.08x
Quality

56%

Does it follow best practices?

Impact

100%

2.08x

Average score across 6 eval scenarios

SecuritybySnyk

Medium

Suggest reviewing before use

Fix and improve this skill with Tessl

tessl review fix ./docs/zh-TW/skills/eval-harness/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

63%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is well-structured and actionable, with executable grader commands, clear templates, and a coherent four-phase workflow, though it lacks an explicit fail-and-retry feedback loop. Its main weaknesses are length from a duplicative worked example and the absence of any progressive disclosure — everything, including reusable templates, is inlined in one ~220-line file.

Suggestions

Move the eval definition templates, grader prompts, and the worked auth example into a references/ file (e.g. references/templates.md) and signal it from SKILL.md, keeping the body to an overview plus workflow

Cut the duplicated 範例 section or shorten it to a pointer, since it restates the full define/implement/evaluate/report workflow already documented

Add an explicit failure-recovery step to the workflow (eval FAIL → fix → re-run eval → only proceed when PASS) to close the feedback loop

DimensionReasoningScore

Conciseness

The body is ~220 lines of mostly templates and process guidance, but it carries redundancy: the full worked example ('範例:新增認證') repeats the define/implement/evaluate/report workflow already documented, and several template blocks restate the same structure twice (workflow section vs. example section). This matches anchor 3 ('mostly efficient but includes some unnecessary explanation or could be tightened'); not 2 since there is no padded conceptual explanation, not 4 given the duplicated example section.

3 / 5

Actionability

Concrete artifacts abound: copy-ready eval definition templates, executable bash checks ('grep -q "export function handleAuth" src/auth.ts && echo "PASS"', 'npm test -- --testPathPattern="auth"'), a file layout, and a full report format. Minor gaps keep it at anchor 4 rather than 5: the '/eval define|check|report' commands are referenced but not defined anywhere in the bundle, and placeholders like '[執行每個能力 eval,記錄 PASS/FAIL]' are pseudocode.

4 / 5

Workflow Clarity

The four-phase workflow (定義 → 實作 → 評估 → 報告) is clearly sequenced with concrete commands and PASS/FAIL checkpoints culminating in explicit status gating ('狀態:準備審查' / '準備發佈'). It sits at anchor 4 ('clear sequence with most checkpoints present; minor validation gaps') rather than 5 because there is no explicit failure-recovery feedback loop (what to do when an eval fails mid-implementation and how to re-run/verify the fix).

4 / 5

Progressive Disclosure

The single file is well-sectioned but everything is inlined: eval templates, grader prompt formats, the metrics reference, and a full worked example — content that belongs in separate reference files (e.g. templates/EXAMPLES.md) at ~220 lines. No bundle files exist to offload it. This fits anchor 3 ('some structure but could be better organized; content that should be separate is inline'); the sectioning is genuinely good, so not 2, but the simple-skill exception does not apply at this length.

3 / 5

Total

14

/

20

Passed

Description

48%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description names a specific niche (eval-driven development) but describes no concrete capabilities and contains no trigger guidance at all, capping completeness at 3. It reads more like a subtitle than an actionable trigger description.

Suggestions

Add 2-3 concrete capability actions, e.g. 'Define capability and regression evals before coding, run them with each change, track pass@k reliability metrics, and generate eval reports'

Add an explicit trigger clause, e.g. 'Use when starting a new feature, when the user asks to write/run evals or test reliability, or before committing changes that risk regressions'

Include natural user phrasings and keywords users would actually say ('evals', 'write evals', 'regression test', 'pass@k', 'EDD') so the skill triggers on real requests

DimensionReasoningScore

Specificity

The phrase "Formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles" names the domain (eval framework, EDD) but lists no concrete actions — nothing like 'define evals', 'run regression checks', 'track pass@k metrics', or 'generate eval reports'. This matches anchor 2 ('Names the domain but actions are minimal or generic'); it is above 1 because the domain is specifically named, and below 3 because no discrete capability is stated.

2 / 5

Completeness

There is a recognizable 'what' ("formal evaluation framework... implementing eval-driven development principles") but zero 'when' guidance — no 'Use when...' clause or equivalent. Per the rubric guideline, a missing 'Use when...' caps completeness at 3; the 'what' itself is also high-level rather than concrete, keeping it at anchor 3 rather than 4.

3 / 5

Trigger Term Quality

Relevant keywords exist ("evaluation", "eval", "EDD", "Claude Code") but natural user phrasings users would actually say — 'write evals', 'run evals', 'test this feature', 'regression check', 'pass@k' — are absent. This fits anchor 3 ('some relevant keywords but missing common variations or synonyms'), not 4, since the natural-term coverage is thin.

3 / 5

Distinctiveness Conflict Risk

The EDD framing ('eval-driven development') carves out a distinct niche separate from general testing or code-review skills, though it could overlap with generic test/CI skills. This matches anchor 4 ('mostly distinct; minor overlap risk with closely related skills'); not 5 because no explicit trigger boundary distinguishes it from adjacent testing skills.

4 / 5

Total

12

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
affaan-m/ECC
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.