CtrlK
BlogDocsLog inGet started
Tessl Logo

tooluniverse-sdk

Build AI scientist systems with the ToolUniverse Python SDK for scientific research. Covers the 3 calling patterns (`tu.run` portable dict API, `tu.tools.X` function API, direct class instantiation), tool loading, batch execution, MCP server integration, and embedding-based tool search. Use for SDK programming, custom tool composition, benchmarking pipelines, and integrating ToolUniverse into research workflows.

68

Quality

84%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

80%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A lean, highly actionable body with copy-paste-ready examples and no padding. It is held back by batch workflows that lack validation checkpoints and by a References section pointing to a REFERENCE.md file that is not present in the bundle.

Suggestions

Add validation checkpoints to the batch/workflow examples: check that targets['data'] exists and inspect run_batch results for per-call errors before dereferencing them, which would lift workflow clarity past the batch-validation cap.

Either add the referenced REFERENCE.md file to the bundle (e.g. under references/) or remove the dangling 'See REFERENCE.md' link so the Resources section only points to real resources.

Include a brief note on how to verify a batch call succeeded (e.g. result shape or error field for run_batch output) in the Batch Execution section.

DimensionReasoningScore

Conciseness

The body is almost entirely executable code and terse reference tables with no explanation of concepts Claude already knows; every section (calling patterns, install, quick start, batch, workflow, config, critical notes, error handling, categories) earns its tokens.

5 / 5

Actionability

The Quick Start and workflow examples are complete and copy-paste ready with real tool names and arguments, and error handling shows concrete exception classes plus a parameter-introspection idiom; the only blemish is an ellipsis placeholder in the error handler, which is too minor to drop a level.

5 / 5

Workflow Clarity

The sequence is present (pattern priority 'start with pattern 1', the 'Always call load_tools()' rule), but the batch workflows (tu.run_batch, the drug-discovery pipeline) include no validation of results — 'targets[\'data\'][:5]' is dereferenced without checking response shape or batch failures, and try/finally is cleanup, not validation. Per the rubric, batch operations without validation cap workflow clarity at 3.

3 / 5

Progressive Disclosure

Sections are well organized and the single reference is clearly signaled one level deep ('See [REFERENCE.md](REFERENCE.md) for detailed guides'), but REFERENCE.md does not exist in the bundle (no references/, scripts/, or assets/ directories), so the sole pointer to deeper detail is dangling and navigation fails.

3 / 5

Total

16

/

20

Passed

Description

88%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: third-person voice, specific and comprehensive about capabilities, with an explicit 'Use for...' clause covering four use contexts. Its only weaknesses are a generic 'SDK programming' trigger and the absence of more natural user-side trigger phrases.

DimensionReasoningScore

Specificity

The description lists six concrete capabilities ('tu.run portable dict API', 'tu.tools.X function API', 'direct class instantiation', 'tool loading', 'batch execution', 'MCP server integration, and embedding-based tool search') with comprehensive coverage of the SDK's surface, matching the 'multiple specific concrete actions; comprehensive coverage' anchor.

5 / 5

Completeness

Both questions are answered explicitly: the what ('Build AI scientist systems with the ToolUniverse Python SDK for scientific research. Covers the 3 calling patterns...') and the when ('Use for SDK programming, custom tool composition, benchmarking pipelines, and integrating ToolUniverse into research workflows').

5 / 5

Trigger Term Quality

'Use for SDK programming, custom tool composition, benchmarking pipelines, and integrating ToolUniverse into research workflows' gives good keyword coverage including the product name, but natural user-side phrases (e.g. querying UniProt/ChEMBL, bioinformatics tool search) are missing and 'SDK programming' is generic.

4 / 5

Distinctiveness Conflict Risk

'ToolUniverse Python SDK' pins a clear niche with distinct triggers, but the generic trigger 'Use for SDK programming' creates minor overlap risk with other SDK-related skills, so it sits just below the 'clear niche with minimal conflict risk' anchor.

4 / 5

Total

18

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 1 missing

Warning

Total

15

/

16

Passed

Repository
mims-harvard/ToolUniverse
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.