CtrlK
BlogDocsLog inGet started
Tessl Logo

evolve

Run an ASI-Evolve style evaluator-driven search workflow for code, algorithms, prompts, or pipelines. Use when Claude needs to align the objective, scoring, evaluator, writable scope, and cognition first, then execute a preflight-gated learn/design/experiment/analyze loop with the bundled Evolve CLI instead of the repository's multi-agent pipeline.

77

Quality

96%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

100%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a dense, executable guide with a well-sequenced workflow, strong validation gates, and clean one-level-deep reference navigation. It respects the token budget while remaining highly actionable.

DimensionReasoningScore

Conciseness

Lean, imperative prose throughout; no padding or re-explanation of concepts Claude already knows — every line is actionable guidance about preserving the operating model, preflight, cognition, and round discipline.

5 / 5

Actionability

Concrete wrapper commands are named (scripts/evolve-brief normalize, scripts/evolve-cognition init/add, evolve-db record/sample/best/stats) alongside an explicit 10-step round loop and specific configuration knobs (sampling.algorithm, island feature semantics, timeout handling).

5 / 5

Workflow Clarity

A clear sequenced round loop with explicit validation/confirmation checkpoints (keep approval.confirmed=false until user confirms; refuse to mutate/evaluate before confirmation; mandatory per-round database sample) plus feedback loops (record lesson, check best snapshot, decide whether another round is justified).

5 / 5

Progressive Disclosure

The body is an overview that points to one-level-deep reference files (operating_model.md, preflight.md, run_spec.md, toolbelt.md, architecture.md) in a clearly signaled 'Read when needed' section; all referenced files exist and the split is appropriate.

5 / 5

Total

20

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific, complete, and distinct, clearly stating both capabilities and trigger conditions. Its only mild weakness is trigger-term breadth, relying more on internal vocabulary than on the full spread of natural user phrasings.

DimensionReasoningScore

Specificity

Names the domain (evaluator-driven search workflow for code/algorithms/prompts/pipelines) and lists multiple concrete actions: aligning objective, scoring, evaluator, writable scope, cognition, then executing a preflight-gated learn/design/experiment/analyze loop.

5 / 5

Completeness

Explicitly answers both 'what' (run an evaluator-driven search workflow with align-then-execute loop) and 'when' (Use when Claude needs to align... then execute the loop), with concrete trigger phrasing.

5 / 5

Trigger Term Quality

Includes natural process terms ('search workflow', 'code, algorithms, prompts, or pipelines', 'learn/design/experiment/analyze loop') and an explicit 'Use when' clause, but leans on internal jargon and omits some everyday synonyms a user might say.

4 / 5

Distinctiveness Conflict Risk

Carves a clear niche by naming the bundled Evolve CLI and explicitly contrasting it with 'the repository's multi-agent pipeline', making conflict with other skills minimal.

5 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
belchman/claude-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.