CtrlK
BlogDocsLog inGet started
Tessl Logo

eval-harness

Formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles

80

2.08x
Quality

51%

Does it follow best practices?

Impact

100%

2.08x

Average score across 6 eval scenarios

SecuritybySnyk

Medium

Suggest reviewing before use

Fix and improve this skill with Tessl

tessl review fix ./docs/zh-TW/skills/eval-harness/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

75%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is concise, actionable, and well-structured with concrete command and template examples and a clear define-implement-evaluate-report workflow. The main gap is the absence of explicit validation/feedback-loop gates in the workflow, which keeps it just short of top marks.

DimensionReasoningScore

Conciseness

The body is largely efficient — short labeled sections, tight code blocks, and minimal preamble; it does not lecture on what evals are, though a few templated markdown blocks (eval/report skeletons) are somewhat padded. Sits above the score-3 'mostly efficient' anchor but short of the fully lean score-5.

4 / 5

Actionability

Provides concrete, runnable shell snippets (grep/npm test/npm run build) and ready-to-use markdown templates for eval definitions and reports; a few template placeholders are non-executable, leaving minor gaps consistent with score-4 'mostly executable, minor gaps'.

4 / 5

Workflow Clarity

A clear four-stage sequence (define → implement → evaluate → report) is laid out with explicit command invocations at each stage; however validation/feedback checkpoints (run eval, fail → fix → re-run) are implied rather than made into explicit validate-then-proceed gates, capping it just below 5.

4 / 5

Progressive Disclosure

Content is well-organized into clearly labeled sections with a navigable structure and no nested references; there are no bundle files, but the body is appropriately self-contained and sectioned, so it lands at score-4 'good structure, minor organization gaps' rather than the fully split score-5.

4 / 5

Total

16

/

20

Passed

Description

28%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description identifies a clear domain but is vague and lacks both concrete actions and an explicit 'Use when' trigger clause. It is closer to the 'Processes PDF files' tier than to the strong examples that list specific actions and triggers.

Suggestions

List concrete actions: 'Define capability and regression evals before coding, run them continuously, and generate pass@k / pass^k reports.'

Add an explicit trigger clause: 'Use when implementing a feature that needs reliability measurement, regression tracking, or eval-driven development.'

Include natural user-facing terms (e.g. 'evals', 'regression checks', 'pass@k') alongside the EDD jargon.

DimensionReasoningScore

Specificity

It names the domain ('eval-driven development') and one abstract action ('implementing EDD principles'), but enumerates no concrete actions like 'define evals', 'run regression checks', or 'generate pass@k reports' — closer to the score-2 anchor than score-3.

2 / 5

Completeness

It states a vague 'what' ('formal evaluation framework') but provides no 'when'/Use-when clause at all; per the guidelines a missing trigger clause caps completeness at 3, and the vague 'what' pulls it to 2.

2 / 5

Trigger Term Quality

Only technical jargon ('eval-driven development', 'EDD', 'eval-driven development (EDD) principles') appears; none of the natural phrases a user would say (e.g. 'write evals', 'check for regressions', 'measure pass@k') are present.

2 / 5

Distinctiveness Conflict Risk

'eval-driven development' is a recognizable niche, but the bare phrasing is generic enough to overlap with testing or general development skills; not clearly distinct, matching the score-3 'somewhat specific but could overlap' anchor.

3 / 5

Total

9

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
affaan-m/ECC
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.