CtrlK
BlogDocsLog inGet started
Tessl Logo

tooluniverse-data-wrangling

Universal data access patterns for downloading and parsing scientific data when ToolUniverse tools don't cover the source, only return metadata, or you need bulk records. Use for VCF/h5ad/BAM/SDF/GCT parsing, multi-step API workflows (search to filter to download to parse), thousands of records at once, or sources with no dedicated tool. Write Python code via Bash for every step.

74

Quality

93%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

The canonical home for this skill is tooluniverse-data-wrangling in mims-harvard/ToolUniverse

SKILL.md
Quality
Evals
Security

Quality

Content

90%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An excellent cookbook-style reference: lean, fully executable code with real endpoints, sensible tool-vs-code decision guidance, and error-handling patterns for batch API work. The two areas short of top marks are the integration of validation checkpoints into the workflows themselves and the amount of content kept inline in SKILL.md despite a working reference-file pattern.

Suggestions

Wire validation into the batch workflows rather than only Section D: e.g., after each bulk-download snippet, add an explicit check that the response is non-empty/valid (status code, HTML-error-page guard, row count) before proceeding to parse.

Apply the existing reference-file pattern to more content: move Section A (format cookbook) or the restricted-sources table into references/ to keep SKILL.md closer to an overview, mirroring how domains 11-24 are handled.

In Section B's per-domain patterns, name the key failure mode to check (e.g., GDC returns 0 hits for over-restrictive filters; ClinicalTrials.gov pageToken loops) so each workflow has a built-in recovery cue.

DimensionReasoningScore

Conciseness

The body is almost entirely dense, executable patterns with one-line comments ("df = pd.read_sas(io.BytesIO(content), format='xport') # SAS Transport (XPT) — NHANES, CDC"); it never explains concepts Claude already knows and adds only non-obvious specifics (URLs, parameters, batching sizes). Every token earns its place, matching the level-5 anchor.

5 / 5

Actionability

Nearly all snippets are copy-paste-ready with real endpoints and parameters, e.g. the NHANES XPT download (requests.get(url) -> pd.read_sas(..., format='xport')) and UniProt cursor pagination loop. Specific examples cover the common cases per format and domain, matching the level-5 anchor; it exceeds level 4 because gaps are essentially absent.

5 / 5

Workflow Clarity

There is a clear decision flow ('When to Use' bullets, 'Decision: Tool vs Code' table, then formats -> domain APIs -> restricted sources -> universal patterns) and Section D supplies validation/retry checkpoints for batch network operations (fetch_with_retry, status-code checks, HTML-error-page guard, raise_for_status). It falls short of level 5 because these checkpoints live in a separate section rather than being wired into each batch workflow as explicit validate-then-proceed steps, leaving some checkpoints implicit.

4 / 5

Progressive Disclosure

The bundle structure is sound: domains 11-24 are deferred to a real one-level-deep reference ([references/specialized-domains.md](references/specialized-domains.md)) that is clearly signaled with a 'When to Read' column per domain. However, the main file inlines ~400 lines of cookbook (Section A, the first ten domain APIs, the restricted-sources table) that could partly live in further reference files, so structure is good but not the clean overview-plus-split of the level-5 anchor.

4 / 5

Total

18

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: concrete actions, explicit and multi-condition 'Use when...' triggers, format-level keywords, and a clearly delineated niche within the ToolUniverse ecosystem. The only gap is a handful of natural trigger terms that were not included, which keeps trigger quality just short of comprehensive.

DimensionReasoningScore

Specificity

The description lists multiple concrete actions with comprehensive coverage: "downloading and parsing scientific data", "VCF/h5ad/BAM/SDF/GCT parsing", "multi-step API workflows (search to filter to download to parse)", and "Write Python code via Bash for every step". It names concrete actions and formats rather than abstract capabilities, matching the level-5 anchor; it exceeds level 4 because coverage is comprehensive rather than having minor gaps.

5 / 5

Completeness

It explicitly answers both questions: the 'what' is "downloading and parsing scientific data ... Write Python code via Bash for every step", and the 'when' is an explicit trigger clause — "when ToolUniverse tools don't cover the source, only return metadata, or you need bulk records" plus "Use for VCF/h5ad/BAM/SDF/GCT parsing, ... or sources with no dedicated tool". This matches the level-5 anchor with concrete trigger phrases and is clearly above level 4's softer 'when'.

5 / 5

Trigger Term Quality

Good keyword coverage with concrete file formats ("VCF/h5ad/BAM/SDF/GCT"), natural phrases like "bulk records", "thousands of records at once", and "sources with no dedicated tool". It sits between level 4 and 5: extension-style format terms and synonyms are present, but a few natural user phrasings (e.g., general data-format terms like FASTA/Excel, or plain 'download data from an API') are absent.

4 / 5

Distinctiveness Conflict Risk

It carves out a clear niche — direct data access when ToolUniverse tools are insufficient — with distinct triggers ("only return metadata", "need bulk records", "no dedicated tool"), so the risk of firing for an unrelated skill is minimal. It is well above level 4 because the ToolUniverse-conditional framing makes it distinguishable even from generic data-analysis skills.

5 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
mims-harvard/ToolUniverse
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.