CtrlK
BlogDocsLog inGet started
Tessl Logo

devtu-self-evolve

Orchestrate the full ToolUniverse self-improvement cycle: discover APIs, create tools, test with researcher personas, fix issues, optimize skills, and push via git. References and dispatches to all other devtu skills. Use when asked to: run the self-improvement loop, do a debug/test round, expand tool coverage, improve tool quality, or evolve ToolUniverse.

69

Quality

85%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

81%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-engineered orchestrator skill: concrete commands, strong validation checkpoints, and real, clearly signaled one-level references. Its main costs are inlined detail (usefulness-testing rubric, issue-category table) that belongs in the existing references and a self-inconsistency between its 150-line target and its actual length.

Suggestions

Move the Skill Usefulness Testing rubric, failure-pattern table, and benchmark commands into a references file (e.g., references/usefulness-testing.md) and keep only the entry-point summary plus a clearly signaled link, restoring the body toward its own 150-line budget.

Replace the Common Issue Categories table with just the existing pointer to references/bug-patterns.md, since the table duplicates content the reference already holds in full.

Inline the persona-agent prompt template (or its key fields) next to the 'Launch 2 agents' step, so the skill's core testing loop is executable without opening the reference.

DimensionReasoningScore

Conciseness

The body is dense and operational (tables, commands, anti-patterns) with no explanation of concepts Claude already knows, but the ~30-line Skill Usefulness Testing rubric/failure table and the Common Issue Categories table (immediately followed by "Full patterns → references/bug-patterns.md") are inlined detail that could be trimmed — minor instances of over-inclusion fitting the 4 anchor, not the padded verbosity of 3.

4 / 5

Actionability

Most phases give copy-paste commands (`python3 -m tooluniverse.cli run <ToolName> '<json_args>'`, `gh pr list --state open`, benchmark script invocations, `ruff check`, git sequences), but the core persona-agent launch describes its parameters rather than giving them verbatim and defers the prompt to the external template — mostly executable with minor gaps (4), not the fully copy-paste-ready 5.

4 / 5

Workflow Clarity

Six clearly sequenced phases with explicit validation checkpoints and feedback loops: mandatory CLI verification before implementing fixes (marked CRITICAL), lint/syntax/run checks in Phase 4, benchmark regression handling ("If any category regresses, prioritize fixing"), the pre-commit fail→re-stage→retry loop, and `"mergeable": "MERGEABLE"` verification before reporting done — matching the 5 anchor's explicit validation and error-recovery loops.

5 / 5

Progressive Disclosure

Both bundle files (references/persona-template.md, references/bug-patterns.md) exist, are linked clearly at point of use, and are one level deep — but the usefulness-testing detail, benchmark commands, and issue-category table are inlined where a 5 would split them out, and the 212-line body violates its own "keep this SKILL.md under 150 lines" rule — good structure with organization gaps (4).

4 / 5

Total

17

/

20

Passed

Description

88%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: third-person, action-enumerated, with an explicit multi-trigger 'Use when' clause. Its only weaknesses are a few missing natural trigger synonyms and inherent overlap with the devtu sub-skills it orchestrates.

DimensionReasoningScore

Specificity

"discover APIs, create tools, test with researcher personas, fix issues, optimize skills, and push via git" lists six concrete, third-person actions covering the full lifecycle — comprehensive coverage per the 5 anchor, with no coverage gaps that would fit the 4 anchor.

5 / 5

Completeness

It explicitly answers both what ("Orchestrate the full ToolUniverse self-improvement cycle: ...") and when ("Use when asked to: run the self-improvement loop, ...") with concrete trigger phrases, matching the 5 anchor exactly; it is above 4 because the 'when' clause is explicit and enumerated rather than merely present.

5 / 5

Trigger Term Quality

The "Use when asked to" list gives natural phrases ("run the self-improvement loop", "do a debug/test round", "expand tool coverage") but misses common variations users would say, such as "test tools" or "add new APIs" — good coverage with a few natural terms missing, matching the 4 anchor rather than the synonym-complete 5.

4 / 5

Distinctiveness Conflict Risk

"evolve ToolUniverse" and the devtu-specific framing establish a clear niche, but because it "dispatches to all other devtu skills," triggers like "improve tool quality" overlap its own sub-skills — minor overlap risk with closely related skills (4), not the minimal-conflict 5.

4 / 5

Total

18

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
mims-harvard/ToolUniverse
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.