CtrlK
BlogDocsLog inGet started
Tessl Logo

eval-harness

克劳德代码会话的正式评估框架,实施评估驱动开发(EDD)原则

44

Quality

44%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Medium

Suggest reviewing before use

Fix and improve this skill with Tessl

tessl review fix ./docs/zh-CN/skills/eval-harness/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

48%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is well-structured and gives usable eval templates and real bash examples, but roughly a quarter of it is duplicated content repeated across the main sections and the 产品评估 (v1.8) section, and its central workflow depends on /eval commands and placeholder steps that are not executable as written. Splitting the worked example and v1.8 material into reference files and adding a failure-feedback loop would materially improve it.

Suggestions

Deduplicate the body: merge the two grader-type lists, the two pass@k/pass^k explanations, and the two eval-storage layouts into single sections — this would cut the file by roughly 25%.

Replace placeholder steps ("[Run each capability eval, record PASS/FAIL]") with executable commands or a bundled script, and either implement the /eval define|check|report commands in a scripts/ file or drop them in favor of concrete file-level instructions.

Move the full worked example and the 产品评估 (v1.8) details into references/ files (e.g. references/example-auth.md, references/product-evals.md) and link to them from SKILL.md, keeping the main file as an overview.

Add an explicit feedback loop to the workflow: after running evals, instruct what to do on FAIL (fix, re-run, only proceed when pass@3 threshold is met).

DimensionReasoningScore

Conciseness

Substantial duplication pads the body: grader types are presented twice ("## 评分器类型" with 3 types, then "### 评分器类型" with 4 in the 产品评估 v1.8 section), pass@k/pass^k guidance appears twice ("## 指标" and "### pass@k 指南"), and the eval file layout is repeated in "## 评估存储" and "### 最小评估工件布局". This matches anchor 2 (several unnecessary/padded sections); not a 3 because the redundancy is structural rather than occasional over-explanation.

2 / 5

Actionability

There are concrete templates (capability/regression eval markdown formats) and real commands ("grep -q \"export function handleAuth\" src/auth.ts && echo PASS", "npm test -- --testPathPattern=\"auth\""), but key execution steps are placeholders ("[Run each capability eval, record PASS/FAIL]") and the /eval define|check|report workflow commands have no backing script or bundle file. This fits anchor 3 (concrete but incomplete, pseudocode in places); not a 4 because central operational steps are not executable as written.

3 / 5

Workflow Clarity

The four phases (定义 → 实现 → 评估 → 报告) are clearly numbered with report formats and status gates ("状态:可以发布", "准备就绪,待审核"), giving explicit checkpoints. It falls short of anchor 5 because there is no fix-and-re-validate feedback loop when an eval fails — the evaluate step records PASS/FAIL but never instructs what to do on failure.

4 / 5

Progressive Disclosure

No references/, scripts/, or assets/ files exist — all ~305 lines live in one SKILL.md, including the full worked example ("## 示例:添加身份验证") and the v1.8 product-eval material that clearly belong in separate reference files. Section headers provide some structure, matching anchor 3; not a 2 because organization exists, not a 4 because content that should be separate is fully inlined.

3 / 5

Total

12

/

20

Passed

Description

41%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description identifies a specific niche (evaluation-driven development for Claude Code sessions) but reads as a category label rather than an actionable trigger: no concrete capabilities are listed and there is no 'use when' guidance. A user needing to define pass/fail criteria or set up pass@k regression tracking would not naturally be matched to it.

Suggestions

Append an explicit trigger clause, e.g. "Use when the user wants to set up evals, define pass/fail criteria for AI-assisted tasks, or measure agent reliability (pass@k)."

List the concrete capabilities instead of the abstract category: defining capability/regression evals, running deterministic or model-based graders, tracking pass@k / pass^k metrics, and generating eval reports.

Add natural synonyms users would actually say ("评估", "回归测试", "基准测试", "eval harness") so the description matches how the need is expressed.

DimensionReasoningScore

Specificity

The description ("克劳德代码会话的正式评估框架,实施评估驱动开发(EDD)原则") names the domain — a formal eval framework for Claude Code sessions — but enumerates no concrete actions such as defining pass/fail criteria, measuring pass@k, or building regression suites. It matches anchor 2 (domain named, actions minimal/generic); it is not a 3 because it never lists even 1-2 discrete actions.

2 / 5

Completeness

A 'what' is present (a formal evaluation framework implementing EDD principles) but the 'when' is entirely absent — there is no 'Use when...' clause or equivalent trigger guidance, which caps completeness at 3 per the judging guidelines. Not a 4 because both halves are not present; not a 2 because the 'what' is clearly stated.

3 / 5

Trigger Term Quality

The only terms are technical jargon ("评估驱动开发(EDD)", "正式评估框架"); there are no natural user phrases like "设置评估", "写回归测试" or file/tool mentions a user would actually say. It sits just above anchor 1 (one or two relevant keywords exist) but well below anchor 3's keyword coverage.

2 / 5

Distinctiveness Conflict Risk

The eval-harness/EDD niche is reasonably specific with distinct vocabulary ("评估驱动开发", "EDD"), giving minor overlap risk only with closely related testing/regression skills. It is not a 5 because without trigger phrases it is harder to cleanly separate from general testing skills.

4 / 5

Total

11

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
affaan-m/ECC
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.