CtrlK
BlogDocsLog inGet started
Tessl Logo

eval-iterate

Iterates on a failing AI/LLM eval (an L2 suite, a golden-set / judge eval, or any eval gating a PR) until it is green AND confirmed, not just luckily passing once. Classifies the failure as a code bug, an eval-definition bug (stale golden item / criteria drift), judge-drift (silent grader version change), or flaky. Applies the minimal fix, then requires N consecutive confirming re-runs — 2 deterministic, 5 for anything a model call grades, since a flat "2" is statistically weak for a stochastic judge. Hard-capped at 5 fix iterations. Refuses to game the eval (Goodhart's Law: skip/delete/overwrite-in-place a case, loosen a threshold) without a second independent check plus `confidence(analysis) >= 90%` and a logged rationale. Composes `ai-engineering`, `confidence`, `critical`, `verify-behavior`. Use when an eval is failing on a PR and needs a real, non-gamed green. Triggers on "this eval is failing", "iterate on this eval", "fix this eval", "get this eval green", "optimize this eval", "/eval-iterate".

72

Quality

91%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

85%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-engineered orchestration skill: concrete commands, an explicit verdict taxonomy, confidence-gated eval edits with hard refusals, and confirmation logic with real feedback loops and a structured exit report. The two real weaknesses are that the entire on-demand loading model depends on rules/ files that are not present in the bundle, and some cross-section repetition of the same gates and rationale.

Suggestions

Ship the three referenced rule files (rules/eval-bug-classification.md, rules/anti-gaming-guard.md, rules/convergence-confirmation.md) in the skill bundle — Phases 2–4 and the anti-gaming gates currently point to files that do not exist, so the load-on-demand design breaks at runtime.

Merge the closing 'Required Reading by Phase' table into the opening phase table — both map phases 2/3/4 to the same rule files, and one table removes the duplication.

State the 2/5 confirmation bar and its rationale once (Phase 4 or convergence-confirmation.md) and reference it elsewhere (e.g. 'N per Phase 4') instead of restating it in Core Principle 1, Phase 4, and the Definition of Done.

DimensionReasoningScore

Conciseness

Operational throughout — commands, verdicts, gates — with no explanations of concepts Claude already knows, but there is trimmable redundancy: the opening phase→rule table and the closing "Required Reading by Phase" table duplicate each other, and the 2/5 confirmation-bar rationale is restated in Phase 4, Core Principle 1, and the Definition of Done. Fits "efficient; minor instances of over-explanation that could be trimmed"; not 3 because none of it is padding or assumed-knowledge explanation.

4 / 5

Actionability

Copy-paste-ready commands throughout — "git rev-parse --abbrev-ref HEAD", "gh pr list --head \"<branch>\" --state open --json number,url --limit 1", "gh pr checks <pr-number> --repo <owner/repo>", "ANTHROPIC_API_KEY=… node scripts/eval/l2.mjs --suite <name>" — plus a complete Phase 6 report template and concrete dispatch strings ("Skill(\"confidence\", \"analysis\")"). Covers the common cases end-to-end; not 4 because there is no gap between instruction and execution.

5 / 5

Workflow Clarity

Phases 0–6 are explicitly sequenced with validation checkpoints at every stage: capture BASELINE_FAILURE "before doing anything else", "Stop at the first failure inside that window", a feedback loop ("return to Phase 2 with the latest failure output as new evidence"), a hard iteration cap with stop conditions, and a Definition of Done checklist. Exemplary match to the top anchor; the destructive-operation cap does not apply because eval edits are gated by independent checks.

5 / 5

Progressive Disclosure

The orchestration-index design is right — a phase→rule table, "Load the matching rule file when you need detail — do not preload them", one-level-deep references — but all three referenced files (rules/eval-bug-classification.md, rules/anti-gaming-guard.md, rules/convergence-confirmation.md) are absent from the bundle, so the disclosure chain breaks at runtime and the detail they promise cannot be verified. Structure alone merits 4, but unresolvable references cap it at 3; not 2 because the in-file organization and signaling are strong, not buried.

3 / 5

Total

17

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An exemplary description: dense with concrete, quantified capabilities, an explicit 'Use when' clause, and six natural trigger phrases including the slash command, all in third person. Its only cost is length — rationale clauses like "since a flat '2' is statistically weak for a stochastic judge" add words beyond bare capability statements — but every clause carries concrete information, so nothing reads as fluff.

DimensionReasoningScore

Specificity

Lists multiple concrete, quantified actions — "Classifies the failure as a code bug, an eval-definition bug (stale golden item / criteria drift), judge-drift … or flaky", "Applies the minimal fix", "requires N consecutive confirming re-runs — 2 deterministic, 5 for anything a model call grades", "Hard-capped at 5 fix iterations" — covering the entire loop with parameters (90% confidence gate, log requirement). Matches the comprehensive-coverage anchor; not 4 because there are no minor gaps in what the skill does.

5 / 5

Completeness

Explicitly answers both what ("Iterates on a failing AI/LLM eval … until it is green AND confirmed", classify → fix → confirm → gate) and when ("Use when an eval is failing on a PR and needs a real, non-gamed green"), followed by concrete trigger phrases. Clear top-anchor match; the ≤3 cap for a missing 'Use when' clause does not apply.

5 / 5

Trigger Term Quality

Six explicit natural trigger phrases — "this eval is failing", "iterate on this eval", "fix this eval", "get this eval green", "optimize this eval", "/eval-iterate" — are exactly what a user would type, reinforced by domain synonyms in the body of the description ("L2 suite", "golden-set / judge eval", "eval gating a PR"). Not 4: synonym coverage of the core noun and its common phrasings is thorough for this domain.

5 / 5

Distinctiveness Conflict Risk

Occupies a clear niche — iterating failing AI/LLM evals with anti-gaming gates — with triggers ("get this eval green", "/eval-iterate") no generic test/debug skill would claim. The named overlap with composed skills ("Composes `ai-engineering`, `confidence`, `critical`, `verify-behavior`") is explicitly framed as composition, keeping conflict risk minimal; not 4 because the trigger surface is distinct rather than merely 'mostly distinct'.

5 / 5

Total

20

/

20

Passed

Validation

75%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 12 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

relative_links

Relative link issues: 16 missing

Warning

referenced_paths_exist

Referenced path issues: 3 missing, 3 deeper-than-1-level

Warning

Total

12

/

16

Passed

Repository
mthines/agent-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.