CtrlK
BlogDocsLog inGet started
Tessl Logo

evolve

Evolve this harness with Darwin Mode — frozen model, evolving harness (real, sandboxed, safety-gated).

56

Quality

64%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./harness/wifi-densepose-sar/.claude/skills/evolve/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

75%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The content is well-structured, actionable, and concise with explicit safety/validation gates and sensible deferral of advanced detail to the external package. Main room for improvement is turning the run flow into an explicit numbered checklist with error-recovery feedback loops.

Suggestions

Add a numbered 'How a generation runs' checklist (mutate one surface file -> sandbox -> run tests -> score -> archive only if improved) with an explicit 'if tests fail, discard' feedback loop to push workflow_clarity toward 5.

Trim the benchmark section to a brief summary and move the full evidence/numbers to the referenced LEARNINGS.md/RESULTS.md to tighten conciseness.

Surface the most common selection/flag combinations inline as one-line examples rather than only deferring them to @metaharness/darwin.

DimensionReasoningScore

Conciseness

The body is lean and assumes Claude's competence; the benchmark section conveys measured, harness-specific empirical findings rather than padding, fitting 'efficient; minor instances of over-explanation that could be trimmed'.

4 / 5

Actionability

Copy-paste-ready commands ('npm run evolve', 'npm run evolve:dry', 'npx metaharness-darwin evolve . --sandbox real --generations 3 --children 4') cover common cases, matching 'mostly executable guidance; concrete commands with minor gaps' on the advanced flags.

4 / 5

Workflow Clarity

The mutate->sandbox->score->archive sequence is clear and validation gates are explicit (validateGeneratedCode, sandbox, tests-pass, measured-improvement), fitting 'clear sequence with most checkpoints present'; not 5 because the user-facing run flow is not a numbered checklist with error-recovery feedback loops.

4 / 5

Progressive Disclosure

Around 50 lines with well-organized sections that defer advanced detail (selection strategies, benchmark evidence) to the external @metaharness/darwin package, matching 'good structure; references mostly clear'; no local bundle files exist to push it to 5.

4 / 5

Total

16

/

20

Passed

Description

53%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description clearly conveys a distinctive niche and concrete qualities but omits an explicit 'when to use' trigger clause, which caps completeness. Adding a 'Use when...' clause with more natural trigger phrases would raise both completeness and trigger_term_quality.

Suggestions

Append an explicit 'Use when...' clause stating when to invoke this skill, e.g. 'Use when you want to self-improve the harness or evolve its surface files.'

Add more natural trigger phrases and synonyms (e.g. 'self-improve', 'mutate the harness', 'optimize the harness') alongside the jargon term 'Darwin Mode'.

List one or two more concrete actions (e.g. 'mutates, sandboxes, scores, and archives harness variants') to lift specificity above 3.

DimensionReasoningScore

Specificity

Names the domain and a couple concrete qualities ('frozen model, evolving harness', 'real, sandboxed, safety-gated') but does not list multiple specific actions, matching the '1-2 concrete actions, not comprehensive' anchor rather than the 'several specific actions' level above.

3 / 5

Completeness

It clearly states what the skill does but provides no 'Use when...' trigger clause, so completeness is capped at 3 per the missing-trigger guidance; it is not 4 because 'when' is absent rather than merely weakly implied.

3 / 5

Trigger Term Quality

Terms like 'evolve', 'Darwin Mode', and 'harness' are relevant to the niche but lean technical and lack common synonyms or natural user phrasings, fitting 'some relevant keywords but missing common variations'.

3 / 5

Distinctiveness Conflict Risk

'Darwin Mode harness evolution' is a clearly distinct niche unlikely to trigger other skills, fitting 'mostly distinct; minor overlap risk'; it is not 5 because the trigger phrasing is somewhat jargon-bound.

4 / 5

Total

13

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
ruvnet/RuView
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.