CtrlK
BlogDocsLog inGet started
Tessl Logo

harness-engineering

This skill should be used when designing autonomous agent harnesses: research loops, evaluation scaffolds, locked and editable surfaces, durable logs, novelty gates, pruning, rollback, PR preparation, and human approval boundaries.

71

Quality

86%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

81%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a well-organized, actionable conceptual skill with explicit workflows, validation, and rollback guidance. Its main gaps are minor discursiveness and references to files outside the skill's bundle.

Suggestions

Trim the discursive framing sentences (e.g., the opening paragraph and 'Autonomy works when feedback is fast...') to tighten conciseness.

Either add the referenced researcher/*.md files to ./references/ or relabel them as external repo paths so progressive disclosure references are verifiable.

Consider splitting the lengthy Detailed Topics / Examples into a separate reference file to reduce the monolithic body while keeping SKILL.md an overview.

DimensionReasoningScore

Conciseness

Largely efficient and well-structured with tables, checklists, and gotchas rather than padding, but the opening paragraph and a few discursive sentences ('Autonomy works when feedback is fast, unambiguous, and hard to game') could be trimmed; above 3 because there is no over-explanation of basic concepts, below 5 due to minor discursiveness.

4 / 5

Actionability

Provides concrete, specific guidance — the surface-class table, an 8-step Harness Design Checklist, a File Layout, and named loop patterns — that is actionable for a design skill; above 3 because the templates are near copy-paste ready, below 5 since the loop diagrams are pattern pseudocode rather than literal commands.

4 / 5

Workflow Clarity

Sequenced loops (autoresearch and research-to-skill) with explicit feedback/rollback ('keep if better -> discard or rollback if worse'), validation steps ('Validate the harness on one known good and one known bad artifact'), and a checklist; not a 4 because checkpoints and error-recovery loops are explicit, and the destructive-operation validation cap is satisfied rather than violated.

5 / 5

Progressive Disclosure

Well-organized into clear sections with a References block that is one level deep and clearly signaled; above 3 because structure and navigation are good, below 5 because the content is monolithic in one file and internal references point to non-bundle paths (researcher/README.md, researcher/rubrics/...) that are not present in ./references/.

4 / 5

Total

17

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is strong: it states a precise niche, lists concrete capabilities comprehensively, and includes an explicit activation clause. Its only weakness is slightly less-than-exhaustive natural trigger synonyms.

DimensionReasoningScore

Specificity

Names the domain ('designing autonomous agent harnesses') and lists multiple concrete capabilities — 'research loops, evaluation scaffolds, locked and editable surfaces, durable logs, novelty gates, pruning, rollback, PR preparation, and human approval boundaries' — giving comprehensive coverage; not below 5 since nothing material is missing, and not a 4 as the list is broad and specific.

5 / 5

Completeness

Explicit 'what' (the enumerated harness surfaces and mechanisms) and an explicit 'when' clause ('This skill should be used when designing...') with concrete triggers; not a 4 because both halves are present and concrete rather than implied.

5 / 5

Trigger Term Quality

Strong domain keywords ('autonomous agent harnesses', 'research loops', 'rollback', 'pruning', 'PR preparation') that target users would naturally say, but a few natural synonyms/variants are absent; above 3 due to good coverage, below 5 because it is not exhaustive.

4 / 5

Distinctiveness Conflict Risk

Occupies a clear niche (autonomous agent control/harness systems) with distinct, specific triggers and minimal overlap with adjacent skills; not a 4 as the framing is sharply scoped to harness design rather than generic agent work.

5 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
muratcankoylan/Agent-Skills-for-Context-Engineering
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.