CtrlK
BlogDocsLog inGet started
Tessl Logo

devtu-create-tool

Create new scientific tools for ToolUniverse framework with proper structure, validation, and testing. Use when users need to add tools to ToolUniverse, implement new API integrations, create tool wrappers for scientific databases/services, expand ToolUniverse capabilities, or follow ToolUniverse contribution guidelines. Supports creating tool classes, JSON configurations, validation, error handling, and test examples.

72

Quality

90%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A highly actionable, well-sequenced skill body built around executable code, explicit validation gates, and a clearly signaled one-level-deep reference structure. The main deductions are minor: duplicated content between the body and testing-guide.md, an unlinked bundle file, and a referenced script (scripts/test_new_tools.py) that is absent from the bundle.

Suggestions

Link references/quick-reference.md from the References section (and drop the duplicate testing-guide link) so every bundle file is discoverable from SKILL.md.

Clarify the location of scripts/test_new_tools.py — either include it in the skill bundle or state that it lives in the ToolUniverse repo — since it is cited as MANDATORY but is not present under ./scripts/.

De-duplicate the verification script: keep the quick commands in SKILL.md and point to references/testing-guide.md for the full script (or vice versa) to reclaim tokens.

DimensionReasoningScore

Conciseness

The body is dense and assumes competence — mostly code, checklists, and terse rules with no basic-concept padding — but has minor trim opportunities: the inline verification script duplicates testing-guide.md nearly verbatim, the testing guide is linked twice, and a time-sensitive claim ("#1 issue in 2026") sits outside any deprecated/old-patterns section. This matches the 'efficient; minor instances that could be trimmed' anchor rather than the every-token-earns-its-place anchor.

4 / 5

Actionability

Fully executable, copy-paste-ready guidance: a complete Python tool class with error handling, a complete JSON tool definition, a runnable 3-step verification script, and four concrete quick commands. Templates like `{API}_{action}_{target}` and concrete limits (55 chars, 150-250 chars, 30s timeout) cover the common cases — a direct match to the top anchor.

5 / 5

Workflow Clarity

The three-step registration sequence is explicit and flagged ("Step 2 MOST COMMONLY MISSED"), testing is a leveled 1-3 checklist with a MANDATORY gate ("test_new_tools.py your_tool -v → 0 failures"), and the verification script asserts each registration step with per-step failure messages. Explicit validation checkpoints plus the 'Top 7 Mistakes' error-avoidance list satisfy the top anchor's validation-and-recovery bar; it is clearly above the 'most checkpoints, minor gaps' anchor.

5 / 5

Progressive Disclosure

Good structure: a tight overview with a clearly signaled, one-level-deep References section pointing to real files (testing-guide.md, advanced-patterns.md, implementation-guide.md, tool-improvement-checklist.md). Two gaps keep it below the top anchor: bundle file references/quick-reference.md is never linked from the body, and the MANDATORY `scripts/test_new_tools.py` is referenced but no scripts/ directory exists in the skill bundle.

4 / 5

Total

18

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: third-person, concrete, and comprehensive, with an explicit 'Use when' clause containing natural trigger phrases tied to the specific ToolUniverse domain. The only minor gap is a few missing synonyms for how users might phrase the request.

DimensionReasoningScore

Specificity

Lists multiple concrete actions — "create tool classes, JSON configurations, validation, error handling, and test examples" plus "implement new API integrations" and "create tool wrappers for scientific databases/services" — covering the skill's capabilities comprehensively; it clearly exceeds the 'several specific actions with minor gaps' anchor.

5 / 5

Completeness

Explicitly answers both: what ("Create new scientific tools ... with proper structure, validation, and testing. Supports creating tool classes, JSON configurations, validation, error handling, and test examples") and when ("Use when users need to add tools to ToolUniverse, implement new API integrations, create tool wrappers ... or follow ToolUniverse contribution guidelines") with concrete trigger phrases — a direct match to the 5 anchor.

5 / 5

Trigger Term Quality

Good natural-phrase coverage: "add tools to ToolUniverse", "implement new API integrations", "create tool wrappers", "follow ToolUniverse contribution guidelines" — phrasings users would plausibly say. A few natural variations (e.g., "register a tool", "integrate an API into ToolUniverse") are missing, so it falls just short of the comprehensive-synonyms anchor.

4 / 5

Distinctiveness Conflict Risk

The ToolUniverse framework name appears in both the capability and trigger clauses, giving it a clear niche with distinct triggers and minimal conflict risk against generic tooling or document skills.

5 / 5

Total

19

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

referenced_paths_exist

Referenced path issues: 2 missing

Warning

Total

15

/

16

Passed

Repository
mims-harvard/ToolUniverse
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.