CtrlK
BlogDocsLog inGet started
Tessl Logo

harness-creator

Build, audit, and improve harnesses that make AI coding agents reliable: AGENTS.md/CLAUDE.md instruction files, feature/state tracking, verification gates, scope boundaries, session handoff, memory persistence, context budgets, tool-permission safety, and multi-agent coordination. Use this whenever a coding agent is unreliable across sessions — forgets context, drifts out of scope, claims "done" before tests pass, or starts each session inconsistently — or when creating or assessing AGENTS.md, CLAUDE.md, feature_list.json, init.sh, progress.md, or session-handoff files. Reach for it even if the user never says the word "harness."

76

Quality

96%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

93%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

Excellent body: lean, executable, and well-structured, with a genuine one-level reference index whose files all exist. The single notable gap is the absence of an explicit post-creation validation/feedback loop connecting create-harness.mjs to validate-harness.mjs.

Suggestions

Add a validation step to the 'Create a harness' workflow: after running create-harness.mjs, run validate-harness.mjs on the target and fix any reported gaps before telling the user the harness is ready.

Note in the audit task what to do when validation surfaces failures (e.g., map the lowest-scoring subsystem back to the corresponding reference pattern), turning the audit into a validate → remediate → re-validate loop.

DimensionReasoningScore

Conciseness

The body is lean throughout: a compact five-row subsystem table, terse "First Move" steps, one-line design rules, and a deliverable checklist. It never explains concepts Claude already knows, and the opening scope line plus "Not for model selection, prompt tuning in isolation..." exclusion earns its place as routing guidance.

5 / 5

Actionability

Every common task ships a copy-paste-ready command with documented flags ("node skills/harness-creator/scripts/create-harness.mjs --target /path/to/project", "--agent-file CLAUDE.md", "--package-manager npm|pnpm|yarn|bun", "--force"). All four referenced scripts exist in the bundle, and the fallback "If you cannot create files, provide exact file contents and commands instead" covers the no-filesystem case.

5 / 5

Workflow Clarity

The "First Move" section gives a clear inspect → ask → minimal-first sequence, destructive operations are gated ("--force only after confirming overwrites are acceptable", "Never hide destructive behavior in scripts"), and the Deliverable Checklist provides an end-state check. However, the create task ends at "explain what was created" with no explicit validate-after-create step or fix-and-retry loop, even though validate-harness.mjs is available for exactly that — a minor validation gap between the 4 and 5 anchors, matching 4.

4 / 5

Progressive Disclosure

"When to Read References" is a textbook one-level-deep reference index: each of the 7 entries names the problem it solves ("Memory across sessions", "Non-obvious failure modes"), and every referenced file (memory-persistence-pattern.md, skill-runtime-pattern.md, tool-registry-pattern.md, context-engineering-pattern.md, multi-agent-pattern.md, lifecycle-bootstrap-pattern.md, gotchas.md) exists in references/. The body stays an overview with details correctly split out.

5 / 5

Total

19

/

20

Passed

Description

96%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: concrete actions, comprehensive natural trigger phrases including exact file names, and explicit what/when structure in third-person/imperative voice with no fluff. The only soft spot is modest overlap risk with CLAUDE.md/AGENTS.md creation skills and a deliberately broadened trigger net.

DimensionReasoningScore

Specificity

The description lists multiple concrete actions — "Build, audit, and improve harnesses" — and comprehensively enumerates the concrete artifacts and subsystems involved ("AGENTS.md/CLAUDE.md instruction files, feature/state tracking, verification gates, scope boundaries, session handoff, memory persistence, context budgets, tool-permission safety, and multi-agent coordination"). Coverage spans the skill's full capability range with no padding.

5 / 5

Completeness

Both 'what' and 'when' are explicitly answered: the opening verb phrase states what it does, and "Use this whenever a coding agent is unreliable across sessions — ... — or when creating or assessing AGENTS.md, CLAUDE.md, feature_list.json, init.sh, progress.md, or session-handoff files" provides concrete trigger phrases for 'when'.

5 / 5

Trigger Term Quality

It includes natural symptom phrases users would actually say ("forgets context, drifts out of scope, claims 'done' before tests pass, or starts each session inconsistently") plus concrete file names users would mention ("AGENTS.md, CLAUDE.md, feature_list.json, init.sh, progress.md, or session-handoff files"). Synonyms and exact file-name triggers are comprehensively covered.

5 / 5

Distinctiveness Conflict Risk

The niche (coding-agent reliability harnesses) and its trigger files (feature_list.json, progress.md, session-handoff) are distinct from most skills, but "creating or assessing AGENTS.md, CLAUDE.md" overlaps with init-style CLAUDE.md-generation skills, and "Reach for it even if the user never says the word 'harness'" deliberately widens the trigger surface. Minor overlap risk with closely related skills rather than the minimal conflict of a fully isolated niche, so it sits between the 4 and 5 anchors, closer to 4.

4 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
walkinglabs/learn-harness-engineering
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.