CtrlK
BlogDocsLog inGet started
Tessl Logo

devtu-benchmark-harness

Continuous improvement system for ToolUniverse tools, skills, and plugin. Run benchmarks, diagnose failures, route fixes to devtu skills, retest. Use after skill optimization, tool additions, or as regression check.

68

Quality

83%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

77%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is an actionable, well-sequenced operations manual with copy-paste commands, explicit validation gates, and retest feedback loops — excellent on actionability and workflow clarity. It loses points on token efficiency (wordy rationale sections and stale question/skill counts) and on progressive disclosure, since references/benchmark-guide.md exists in the bundle but is never referenced from the body.

Suggestions

Link references/benchmark-guide.md from the body (e.g., under 'Available benchmarks': 'Per-benchmark breakdowns, coverage, and score history: see references/benchmark-guide.md') so the sole reference file is discoverable, and reconcile its stale question counts (61 vs the body's 205) with the body's tables.

Move volatile, time-sensitive numbers (question counts like 205/20, sub-skill count 113, current best scores) into benchmark-guide.md or a tracking file, keeping only structural facts in SKILL.md so the body does not drift.

Tighten the APPEND_CONVENTIONS section: replace the multi-sentence rationale with one line ('APPEND_CONVENTIONS=1 forces conventions into every system prompt — measures convention correctness; default mode measures routing reliability; the gap is the routing problem') and trim the grader strategy list to a pointer at grade_answers.py.

DimensionReasoningScore

Conciseness

The body is mostly efficient and command-first, but the APPEND_CONVENTIONS rationale paragraph and the plugin-architecture explanation are wordy, and it embeds drift-prone specifics ("205 computational", "113 sub-skills", "20 MCQ") that will go stale — time-sensitive numbers not quarantined in a deprecated section. It is not 4 because the trimming needed goes beyond minor instances, and not 2 because the bulk is genuinely operational rather than padded.

3 / 5

Actionability

Fully executable copy-paste commands with flags throughout (run_harness_loop.sh, check_memorization.py, run_eval.py invocations), plus a worked diagnosis example showing the exact message to hand the devtu skill. Anchor 5: specific commands and examples cover the common cases.

5 / 5

Workflow Clarity

The 5-step loop is explicitly sequenced with an orchestrated one-command runner, a validation gate ("Before accepting any skill edit ... run: check_memorization.py --all"), a root-cause verification procedure ("Run the script yourself — does it reproduce the GT value?"), and a retest feedback loop ("Compare: how many flipped from wrong to correct?"). Anchor 5: explicit validation steps and feedback loops for error recovery.

5 / 5

Progressive Disclosure

Scripts are properly referenced one level deep and the body is well-sectioned, but the bundle's only references file (references/benchmark-guide.md) is never mentioned in the body, making it undiscoverable, and bulky detail (7 grader strategies, plugin architecture rationale, failure-pattern tables) is inlined where that reference file would be the natural home — the guide even duplicates body content with conflicting stale counts (61 vs 205 questions). Anchor 3 fits; not 4 because an entire bundle reference is un-signaled, which is more than a minor organization gap.

3 / 5

Total

16

/

20

Passed

Description

88%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: third-person, concise, concrete about actions, with an explicit 'Use after...' trigger clause covering three situations. Its only weaknesses are missing natural synonyms (eval/evaluation/score) and mild overlap risk with the devtu skills it dispatches to.

DimensionReasoningScore

Specificity

"Run benchmarks, diagnose failures, route fixes to devtu skills, retest" lists four distinct concrete actions that comprehensively cover the improvement loop, matching the anchor-5 example's shape. It is not 4 because there is no meaningful coverage gap in the action list.

5 / 5

Completeness

The description explicitly answers both questions: what ("Run benchmarks, diagnose failures, route fixes to devtu skills, retest") and when ("Use after skill optimization, tool additions, or as regression check") with three concrete trigger situations. It is not 4 because the when-clause is explicit and specific rather than merely present.

5 / 5

Trigger Term Quality

Natural terms like "benchmarks", "failures", "retest", "skill optimization", "tool additions", and "regression check" give good keyword coverage a user would plausibly say. It falls short of 5 because common synonyms and variations ("eval", "evaluation", "scores", "accuracy") are missing.

4 / 5

Distinctiveness Conflict Risk

"Continuous improvement system for ToolUniverse tools, skills, and plugin" is clearly bound to a niche ecosystem, but it overlaps with the closely related devtu-fix-tool / devtu-optimize-skills / devtu-self-evolve skills it routes to — a user asking to fix a tool could match this harness instead. Mostly distinct with minor overlap risk, matching anchor 4; not 5 because that routing overlap is real.

4 / 5

Total

18

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

referenced_paths_exist

Referenced path issues: 2 missing

Warning

Total

15

/

16

Passed

Repository
mims-harvard/ToolUniverse
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.