CtrlK
BlogDocsLog inGet started
Tessl Logo

eval-harness

Claude Codeセッションの正式な評価フレームワークで、評価駆動開発(EDD)の原則を実装します

51

Quality

56%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Medium

Suggest reviewing before use

Fix and improve this skill with Tessl

tessl review fix ./docs/ja-JP/skills/eval-harness/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

75%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The content is a well-structured, concrete guide: it provides executable grader commands, copy-paste eval templates, a clear define-implement-evaluate-report workflow, and framework-specific metrics. Its main weaknesses are minor padding, a few placeholders that aren't fully executable, and no reference-file split despite the length.

DimensionReasoningScore

Conciseness

The body is mostly templates, code blocks, and framework-specific metric definitions (pass@k/pass^k) rather than explanations of concepts Claude already knows; minor padding exists in the repeated report templates and the philosophy section.

4 / 5

Actionability

Mostly executable guidance: real grader commands (grep -q, npm test, npm run build), copy-paste markdown eval templates, and /eval commands. Minor gaps are the placeholder '[各能力評価を実行し、PASS/FAILを記録]' and /eval slash-commands with no backing implementation.

4 / 5

Workflow Clarity

A clear four-phase sequence (定義 → 実装 → 評価 → レポート) with per-eval PASS/FAIL recording and report formats; validation checkpoints exist (running regression evals after implementation) but explicit validate→fix→retry feedback loops are only implicit.

4 / 5

Progressive Disclosure

No bundle files exist, and the single body is well-organized with clear section headers; however, at ~220 lines some content (the worked auth example, evaluator type details) could be split into reference files for easier navigation.

4 / 5

Total

16

/

20

Passed

Description

37%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description states a clear niche (evaluation-driven development for Claude Code sessions) but is too vague to act as a trigger: it names no concrete actions, includes no 'use when' guidance, and lacks the natural keywords a user would actually say. It would rarely be surfaced when needed.

Suggestions

List concrete actions in the description, e.g. 'Defines capability and regression evals before coding, runs them via code/model/human graders, and tracks pass@k reliability metrics.'

Add an explicit trigger clause, e.g. 'Use when starting a new feature, defining success criteria before implementation, or checking whether changes broke existing functionality.'

Include natural trigger terms users would say — 'eval', 'evaluate', 'regression test', 'pass@k', 'success criteria', 'test reliability' — to improve retrieval.

DimensionReasoningScore

Specificity

The description names the domain ('Claude Codeセッションの評価フレームワーク') but its only stated action is the abstract '評価駆動開発(EDD)の原則を実装します' — no concrete verbs like defining evals, running regression checks, or tracking pass@k.

2 / 5

Completeness

A 'what' is present (an evaluation framework implementing EDD principles) but there is no 'when' or 'Use when...' trigger clause anywhere, which caps completeness at 3 per the rubric guidelines.

3 / 5

Trigger Term Quality

Keywords are limited to 'Claude Code', '評価フレームワーク', and 'EDD'; natural user phrases such as 'eval', 'regression test', 'pass@k', or 'test before coding' are entirely absent.

2 / 5

Distinctiveness Conflict Risk

The Claude Code session + EDD niche is somewhat specific, but the generic 'evaluation framework' framing overlaps with test/review/verify skills and no distinct trigger phrases separate it.

3 / 5

Total

10

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
affaan-m/ECC
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.