CtrlK
BlogDocsLog inGet started
Tessl Logo

harness-eval

Evaluate a repo agent harness (AGENTS.md, rules, skills, skill refs) for broken paths/commands, redundant instructions, and usefulness using a stack-agnostic dual-judge protocol with planted traps. HIGH PRIORITY questionnaires at top: Q1 optional docs, Q2 B/C budget before Track A (certainty/tokens). A always runs after Q2; B/C opt-in. ADRs/RFCs excluded from T2. Mixed apply uses 11-mixed-apply.md (KEEP/CUT). Use when the user says harness eval, harness-eval, harness debug, audit AGENTS.md, audit skills/rules, instruction audit, redundancy of agent instructions, usefulness of skills, Ship/Review/Hold/Slim/Keep-core for harness, or wants Track A/B/C harness evaluation. Do NOT use for harness setup or init, feature spec-driven work (tlc-spec-driven), or applying Ship/Slim trims unless the user explicitly asks after the report.

77

Quality

97%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

92%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a well-structured, actionable protocol: sequenced steps with concrete commands, explicit validation/feedback loops for destructive operations, and clean one-level-deep references to real bundle files. Its only weakness is mild verbosity and restatement between the Critical rules, Steps, and Troubleshooting sections.

DimensionReasoningScore

Conciseness

The body is dense but justified by a genuinely complex multi-track protocol; it assumes Claude's competence (concrete commands and rules, no 'what is a PDF' filler). It is a 4 rather than 5 because some Critical rules and Troubleshooting entries restate protocol mechanics that could be tightened, and the 14-rule list carries minor redundancy with the step instructions.

4 / 5

Actionability

Provides copy-paste-ready, executable bash commands for every step (e.g. `python3 "$SKILL_DIR/scripts/inventory_extract.py" --root . --run-id "$RUN_ID"`) and references scripts that all exist in ./scripts/; not a 4 because the guidance is fully executable and covers the common cases with concrete flags and expected outputs.

5 / 5

Workflow Clarity

Steps 1–11 are clearly sequenced with explicit validation checkpoints and feedback loops (trap gate PASS/FAIL, 'If missing, the skill install is broken — stop', re-merge on trap FAIL, fan-in PASS before Slim). Despite destructive Slim deletes, validation is present, so the destructive-cap does not apply; not a 4 because checkpoints and error-recovery loops are explicit rather than minor-gapped.

5 / 5

Progressive Disclosure

Self-contained skill with a clear overview body pointing one level deep to real references (PROTOCOL.md, judge-prompts.md, GLOSSARY.md, claims.schema.json) and scripts, all verified present in ./references/ and ./scripts/; navigation is explicitly signaled ('Read references/PROTOCOL.md completely before the first run'). Not a 4 because the split is appropriate and references are clearly signaled rather than minorly disorganized.

5 / 5

Total

19

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

This is a strong, third-person description that explicitly covers what the skill does, when to use it, and when not to. Trigger terms are comprehensive and natural, and the negative-scope clause sharply reduces conflict risk. It is dense with procedural internals (Q1/Q2, ADR handling, Mixed apply) that lean toward overload, but the rubric scores explicit content rather than verbosity.

DimensionReasoningScore

Specificity

Names the domain (repo agent harness: AGENTS.md, rules, skills, skill refs) and multiple concrete actions — 'broken paths/commands', 'redundant instructions', and 'usefulness' — matching the comprehensive-coverage anchor; not a 4 because the action set is explicit and complete rather than having minor gaps.

5 / 5

Completeness

Explicitly answers both 'what' ('Evaluate a repo agent harness... for broken paths/commands, redundant instructions, and usefulness using a stack-agnostic dual-judge protocol with planted traps') and 'when' ('Use when the user says...') with concrete trigger phrases; not a 4 because both halves are explicit and specific rather than the 'when' being merely present.

5 / 5

Trigger Term Quality

Lists natural trigger phrases users would actually say — 'harness eval', 'harness-eval', 'harness debug', 'audit AGENTS.md', 'audit skills/rules', 'instruction audit', 'redundancy of agent instructions', 'usefulness of skills', 'Ship/Review/Hold/Slim/Keep-core', 'Track A/B/C harness evaluation' — comprehensive coverage with synonyms; not a 4 because no common natural variants are missing.

5 / 5

Distinctiveness Conflict Risk

Clear niche (harness evaluation) with distinct triggers and explicit negative scope ('Do NOT use for harness setup or init, feature spec-driven work (tlc-spec-driven), or applying Ship/Slim trims unless...'), minimizing conflict risk; not a 4 because the negative boundary removes the residual overlap risk.

5 / 5

Total

20

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
tech-leads-club/agent-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.