CtrlK
BlogDocsLog inGet started
Tessl Logo

evolve

Evolve this harness with Darwin Mode — frozen model, evolving harness (real, sandboxed, safety-gated).

54

Quality

60%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./harnesses/timesfm-harness/.claude/skills/evolve/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

78%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a lean, well-structured overview with concrete executable commands and genuine safety/validation gates baked into the described evolution loop. It assumes Claude's intelligence and defers deep detail to external package docs, scoring strongly across all dimensions.

DimensionReasoningScore

Conciseness

The body is dense and information-rich with no basic-concept padding; the benchmarks section is lengthy but consists of measured insights rather than filler, with only minor trims possible. Time-sensitive stats (7.7% -> 15.3%, 31x cheaper) are confined to a dedicated 'What the benchmarks taught us' section that defers full evidence to LEARNINGS.md, so they do not unduly penalize conciseness.

4 / 5

Actionability

It provides copy-paste-ready commands (`npm run evolve`, `npm run evolve:dry`, the `npx` invocation with concrete flags) covering the common cases, with only minor gaps around the advanced selection/statistical flags being mentioned but not demonstrated.

4 / 5

Workflow Clarity

The internal loop (mutate one surface file -> sandbox -> score -> keep only measurably-improving variants) is a clear sequence backed by explicit validation checkpoints (validateGeneratedCode gate, tests must pass, measured-improvement promotion guard), with only minor gaps in presenting it as an explicit numbered user workflow.

4 / 5

Progressive Disclosure

The body is a concise overview (~40 lines) organized into clear sections (Run it, Safety, benchmarks) with well-signaled one-level-deep references to `@metaharness/darwin` and its LEARNINGS.md / bench/results/RESULTS.md for advanced detail; no local bundle files are needed.

5 / 5

Total

17

/

20

Passed

Description

41%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description communicates a specific, distinctive niche but is written in imperative voice, leans on technical jargon over natural trigger terms, and omits any explicit "Use when..." guidance. It is recognizable but would benefit from third-person phrasing and concrete trigger conditions.

Suggestions

Rewrite in third person (e.g. "Evolved the harness via Darwin Mode...") and add a concrete "Use when the user wants to improve the harness automatically / run Darwin Mode" trigger clause.

Swap jargon for natural keywords users would say, such as "self-improve the harness", "auto-tune the harness", or "run Darwin Mode evolution".

List 1-2 concrete actions (mutate surface files, sandbox and score variants, archive improving descendants) to lift specificity above the single verb "Evolve".

DimensionReasoningScore

Specificity

The description names the domain and a couple of concrete characteristics ("sandboxed, safety-gated") but offers only the single generic action "Evolve"; the imperative "Evolve this harness" is second-person voice, which the rubric penalizes, pulling it down from a borderline 3.

2 / 5

Completeness

It gives a clear "what" (evolve the harness via Darwin Mode with a frozen model) but includes no "Use when..." clause or equivalent trigger guidance, so completeness is capped at 3 per the rubric guideline.

3 / 5

Trigger Term Quality

Terms like "Darwin Mode", "frozen model, evolving harness" are niche technical jargon; only "evolve" and "harness" approach natural phrasing, so most common trigger phrases a user would actually say are missing.

2 / 5

Distinctiveness Conflict Risk

It carves a clear niche — harness self-evolution with Darwin Mode — that is unlikely to fire for unrelated skills, with only minor overlap risk against other harness/tooling skills.

4 / 5

Total

11

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
ruvnet/RuVector
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.