CtrlK
BlogDocsLog inGet started
Tessl Logo

tooluniverse-sdk

Build AI scientist systems with the ToolUniverse Python SDK for scientific research. Covers the 3 calling patterns (`tu.run` portable dict API, `tu.tools.X` function API, direct class instantiation), tool loading, batch execution, MCP server integration, and embedding-based tool search. Use for SDK programming, custom tool composition, benchmarking pipelines, and integrating ToolUniverse into research workflows.

70

Quality

86%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

The canonical home for this skill is tooluniverse-sdk in mims-harvard/ToolUniverse

SKILL.md
Quality
Evals
Security

Quality

Content

80%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is an efficient, highly actionable SDK reference with excellent gotcha coverage, but it stumbles on the two structural dimensions: batch workflows lack validation/verification feedback loops (capping workflow clarity), and its only progressive-disclosure pointer targets a REFERENCE.md that does not exist in the bundle.

Suggestions

Add a validation/verification step around batch execution, e.g. check each result in run_batch output for error structures before proceeding, and show a fix-and-retry loop in the drug-discovery pipeline example.

Either ship the referenced REFERENCE.md (moving the tool-categories table and detailed per-category API usage into it) or remove the dangling link so navigation is not a dead end.

Move the Tool Categories table into the reference layer and keep SKILL.md to the calling patterns, quick start, and critical notes to sharpen the overview/navigation split.

DimensionReasoningScore

Conciseness

The body is lean and entirely SDK-specific — no explanation of concepts Claude already knows (no "what is an API" padding), minimal purposeful comments ("REQUIRED before any tool call"), and dense gotchas like case-sensitive tool names and nested tool-finder output. Every section earns its tokens; this matches the anchor 'lean and efficient; assumes Claude's competence'.

5 / 5

Actionability

Code is copy-paste ready and covers the common cases: install variants and env vars, load_tools-then-call quick start with real tool names and arguments ("UniProt_get_entry_by_accession", accession "P05067"), run_batch, caching config, and a typed exception-handling block. This matches the anchor 'fully executable; copy-paste ready; specific examples cover the common cases'.

5 / 5

Workflow Clarity

A sequence is present (install → load_tools → find tools → execute → batch), and Critical Notes include useful pre-checks (isinstance check, required-params lookup), but the batch execution and drug-discovery pipeline workflows have no validation/verification of results and no fix-and-retry loop — the try/finally is cleanup, not checkpointing. Per the rubric, batch operations without validation cap workflow clarity at 3, which fits the anchor 'steps listed but validation gaps'.

3 / 5

Progressive Disclosure

The single external pointer ("See [REFERENCE.md](REFERENCE.md) for detailed guides") is one level deep and clearly signaled, but no references/, scripts/, or assets/ directories exist and REFERENCE.md is absent from the bundle — the link is dangling. At ~130 lines with the tool-categories table and error-handling detail inlined, this matches the anchor 'some structure... content that should be separate is inline' rather than the good-organization anchor above it.

3 / 5

Total

16

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: third-person, specific to the SDK's real API patterns, with an explicit and concrete "Use for" trigger clause. The only weakness is trigger-term coverage that omits some natural synonyms a user might say when they need this skill.

DimensionReasoningScore

Specificity

The description lists multiple concrete actions with actual API surface: "the 3 calling patterns (`tu.run` portable dict API, `tu.tools.X` function API, direct class instantiation), tool loading, batch execution, MCP server integration, and embedding-based tool search". This matches the anchor for comprehensive, specific concrete actions; nothing here is vague filler.

5 / 5

Completeness

It explicitly answers both questions: the "what" via the enumerated capabilities (calling patterns, tool loading, batch execution, MCP integration, embedding search) and the "when" via a concrete "Use for..." clause with four trigger scenarios. This mirrors the top anchor's structure of explicit what-and-when with concrete triggers.

5 / 5

Trigger Term Quality

"Use for SDK programming, custom tool composition, benchmarking pipelines, and integrating ToolUniverse into research workflows" gives good keyword coverage, but a few natural user phrasings are missing (e.g. "calling scientific APIs/tools", "AI scientist agent", "drug discovery tools", specific bioinformatics services). Not score 5 because synonym-level coverage of what a user would naturally say is incomplete; not score 3 because several genuinely natural trigger phrases are present.

4 / 5

Distinctiveness Conflict Risk

"Build AI scientist systems with the ToolUniverse Python SDK" carves out a clear niche around a named SDK, and triggers are tied to that SDK and its workflows, so conflict with other skills is minimal. It is not merely broad domain language like "helps with research code".

5 / 5

Total

19

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 1 missing

Warning

Total

15

/

16

Passed

Repository
mims-harvard/ToolUniverse
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.