CtrlK
BlogDocsLog inGet started
Tessl Logo

devtu-fix-tool

Fix failing ToolUniverse tools by diagnosing test failures, identifying root causes, implementing fixes, and validating solutions. Use when ToolUniverse tools fail tests, return errors, have schema validation issues, or when asked to debug or fix tools in the ToolUniverse framework.

65

Quality

79%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/devtu-fix-tool/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

71%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A highly actionable, domain-expert debugging guide whose concrete commands, symptom→fix mapping, and validation checkpoints are excellent. Its weaknesses are redundancy (duplicated tables and repeated verification sections), a step-numbering error, and a monolithic structure with a dangling reference to a nonexistent unit-tests-reference.md file.

Suggestions

Deduplicate the content: merge the 'Where to Fix' and 'Quick Reference' error tables into one, and fold the 'Verification', 'Unit Test Management', and 'Common Pitfalls' recaps into the main Instructions workflow to remove repeated material.

Fix the broken [unit-tests-reference.md](unit-tests-reference.md) reference — no such file exists in the bundle; either create it with the detailed unit-test patterns or remove the link and keep only the inline checklist.

Move the 10-type 'Error Types' catalog (and detailed unit-test examples) into a separate reference file so SKILL.md stays a lean overview, and renumber the Instructions list (steps 4 is duplicated) so the sequence reads 1–7.

DimensionReasoningScore

Conciseness

The body is dense with domain-specific knowledge (symptom/cause/fix tuples, exact commands) but contains several redundant sections: the 'Where to Fix' and 'Quick Reference' error tables repeat each other, 'Verification' and 'Unit Test Management' restate the Instructions steps, and 'Common Pitfalls' recaps earlier content. Mostly efficient, but clearly could be tightened — not the 'minor instances' level of a 4.

3 / 5

Actionability

Fully executable guidance throughout: copy-paste commands ('python scripts/test_new_tools.py <pattern> -v', 'python -m tooluniverse.generate_tools'), exact JSON config snippets ("{\"type\": [\"integer\", \"null\"]}"), per-error-type file-path tables, and concrete wrong-vs-right code examples. Not a 4 because commands and examples are complete and cover the common cases rather than having gaps.

5 / 5

Workflow Clarity

The Instructions give a clear diagnose→fix→regenerate→test sequence with explicit validation checkpoints (verify the bug via CLI first, re-run integration and unit tests, a unit-test checklist). Held below 5 because the numbered list contains two steps numbered '4.' and the validation steps are restated across three separate sections, blurring the canonical sequence.

4 / 5

Progressive Disclosure

Good section headers and tables, but the ~380-line body is monolithic: the 10-type error catalog and unit-test patterns are inlined rather than split into reference files. The sole reference, [unit-tests-reference.md](unit-tests-reference.md), points to a file that does not exist in the bundle (no references/ directory), which undermines navigation. Not a 2 because the in-body structure itself is well organized.

3 / 5

Total

15

/

20

Passed

Description

87%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description with a clear what/when structure, third-person imperative voice, and a well-scoped niche that minimizes conflict risk. Its only weakness is that the capability list uses generic debugging verbs while the concrete error classes live in the trigger clause, leaving minor specificity and trigger-synonym gaps.

DimensionReasoningScore

Specificity

Quotes four concrete actions — 'diagnosing test failures, identifying root causes, implementing fixes, and validating solutions' — covering the full fix lifecycle. Below the comprehensive anchor (5) because the action verbs are process-generic debugging language; the specific error classes (schema validation, JSON parsing, parameter errors) appear only in the trigger clause, leaving minor coverage gaps.

4 / 5

Completeness

Explicitly answers both questions: the 'what' ('diagnosing test failures, identifying root causes, implementing fixes, and validating solutions') and a concrete multi-trigger 'when' ('Use when ToolUniverse tools fail tests, return errors, have schema validation issues, or when asked to debug or fix tools'). This matches the top anchor's pattern of what + explicit trigger phrases.

5 / 5

Trigger Term Quality

Natural user phrases are present: 'fail tests', 'return errors', 'schema validation issues', 'debug or fix tools', all anchored to the ToolUniverse domain. Good coverage, but common variations like 'broken tools', 'test failures', or 'failing tests' as standalone triggers are missing, so it falls short of the synonym-complete anchor (5).

4 / 5

Distinctiveness Conflict Risk

The niche is unambiguous: 'ToolUniverse' is named in both sentences and every trigger is framework-specific, so it would not fire for generic debugging or testing requests. Minimal conflict risk, matching the clear-niche anchor (5).

5 / 5

Total

18

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 1 missing

Warning

referenced_paths_exist

Referenced path issues: 3 missing

Warning

Total

14

/

16

Passed

Repository
mims-harvard/ToolUniverse
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.