CtrlK
BlogDocsLog inGet started
Tessl Logo

eval-harness-first

Build the evaluation harness that gates every fine-tuning run — golden sets, per-failure-mode graders, judge calibration, and base-model baselines. Use when starting a fine-tuning effort, when converting traces into an eval set, or when calibrating a judge against human labels.

72

Quality

88%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

85%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured process skill: clear sequenced workflow with validation checkpoints and feedback loops, and clean one-level-deep references to real bundle files. Minor conciseness redundancy and the absence of runnable code keep actionability and conciseness just below the top anchor.

Suggestions

State the all-deterministic judge-calibration N/A path once (in the Judge Calibration section) and have the Exit Checklist reference it, instead of restating the condition in three places.

Add a short runnable snippet for one grader shape inline (e.g. a schema-compliance pass/fail) so the Graders section shows executable guidance, not only pointers.

Tighten the Building Goldens bullets by collapsing the synthetic-generation rationale, which slightly overlaps the dimension-based-generation description.

DimensionReasoningScore

Conciseness

The body is lean and assumes Claude's competence without explaining what fine-tuning is, but the all-deterministic N/A path is restated in three places (Building Goldens, Judge Calibration, Exit Checklist), a minor redundancy that could be trimmed.

4 / 5

Actionability

Concrete thresholds (≥100 traces, 4–8 buckets, TPR/TNR), a precise directory contract, and exact file paths give mostly executable guidance; as an instruction-only skill it lacks runnable code but the actionable specifics largely compensate.

4 / 5

Workflow Clarity

The 8-step flywheel is clearly sequenced, the Phase 0 Exit Checklist is an explicit validation gate, and drift detection feeds back to step 2 — a full feedback loop with a checklist for a batch/training process.

5 / 5

Progressive Disclosure

Overview points to two real, one-level-deep reference files (grader-templates.md, judge-calibration.md), each clearly signaled inline with what it contains; both files exist in the bundle.

5 / 5

Total

18

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, specific description with clear what/when structure and concrete trigger phrases in third-person voice. Minor keyword coverage gaps keep trigger quality just below the top anchor.

DimensionReasoningScore

Specificity

"Build the evaluation harness... golden sets, per-failure-mode graders, judge calibration, and base-model baselines" names multiple distinct concrete deliverables with comprehensive coverage of the harness's components.

5 / 5

Completeness

It explicitly answers what ("Build the evaluation harness...golden sets, graders, calibration, baselines") and when ("Use when starting a fine-tuning effort, when converting traces... or when calibrating a judge") with concrete trigger phrases.

5 / 5

Trigger Term Quality

"fine-tuning effort", "converting traces into an eval set", "calibrating a judge against human labels" are natural domain phrases, but a few common synonyms (e.g. 'evals', 'benchmarking') are absent.

4 / 5

Distinctiveness Conflict Risk

The fine-tuning eval-harness niche has distinct, specific triggers (golden sets, per-failure-mode graders, judge calibration) with minimal overlap risk against other skills.

5 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
wshobson/agents
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.