CtrlK
BlogDocsLog inGet started
Tessl Logo

tooluniverse-data-wrangling

Universal data access patterns for downloading and parsing scientific data when ToolUniverse tools don't cover the source, only return metadata, or you need bulk records. Use for VCF/h5ad/BAM/SDF/GCT parsing, multi-step API workflows (search to filter to download to parse), thousands of records at once, or sources with no dedicated tool. Write Python code via Bash for every step.

71

Quality

89%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

86%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An excellent reference-style skill: fully executable code, clear tool-vs-code decision guidance, and a well-structured one-level-deep bundle split for specialized domains. The only weaknesses are minor duplication of pagination patterns between Sections B and D and the absence of an explicit validation checkpoint after bulk fetches.

Suggestions

Remove the per-API pagination loops in Section B (#2, #8, #9) or reduce them to one-line pointers to the generic pagination patterns in Section D to cut duplicated code.

Add an explicit post-batch validation checkpoint, e.g., after a bulk fetch assert len(all_records) matches the reported hitCount/total before parsing, since the skill targets batch operations with thousands of records.

Consider offloading the Section C restricted-sources table into the existing reference file to further slim the main SKILL.md, keeping only the most common restricted sources inline.

DimensionReasoningScore

Conciseness

The body is a dense, prose-free cookbook that assumes Claude's competence — every line is a runnable pattern or a table row. Not 5: pagination appears both in Section B examples (#2 UniProt, #8 ClinicalTrials.gov, #9 EuropePMC) and again generically in Section D, a duplication that could be trimmed.

4 / 5

Actionability

Fully executable copy-paste code throughout: real API endpoints (eutils, GDC, GTEx storage), working filter payloads, complete batch-fetch loops, and a runnable retry helper. Not 4: there are no pseudocode gaps; the examples cover the common cases end-to-end.

5 / 5

Workflow Clarity

A decision table (Tool vs Code), explicit search->filter->download->parse sequences, and an error-handling section with status-code checks and an HTML-error-page guard provide most checkpoints. Not 5: batch/bulk workflows lack an explicit post-fetch validation step (e.g., verify expected record count before parsing).

4 / 5

Progressive Disclosure

SKILL.md keeps core patterns inline and offloads 14 specialized domains to a real, one-level-deep reference (references/specialized-domains.md), clearly signaled via 'read references/specialized-domains.md when you need...' plus a When-to-Read summary table. Not 4: the split is appropriate and navigation is easy, with no buried or nested references.

5 / 5

Total

18

/

20

Passed

Description

88%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description with a clear what/when structure, concrete actions, and domain-natural trigger terms anchored to the ToolUniverse-gap niche. Minor gaps in synonym coverage and slight overlap risk with generic data-download skills keep two dimensions at 4.

DimensionReasoningScore

Specificity

Concrete actions are named comprehensively — 'downloading and parsing scientific data', 'multi-step API workflows (search to filter to download to parse)', 'thousands of records at once', 'Write Python code via Bash' — plus specific formats (VCF/h5ad/BAM/SDF/GCT). Not 4: action coverage is comprehensive rather than having minor gaps.

5 / 5

Completeness

Explicitly answers both: what ('Universal data access patterns for downloading and parsing scientific data') and when ('Use for VCF/h5ad/BAM/SDF/GCT parsing... thousands of records at once, or sources with no dedicated tool') with concrete trigger phrases. Not 4: the 'when' is fully explicit, not merely present.

5 / 5

Trigger Term Quality

Natural domain terms like 'VCF/h5ad/BAM/SDF/GCT parsing', 'bulk records', and 'API workflows' are present and would be said by target users. Not 5: some natural synonyms are missing (e.g., 'download a dataset', other common formats like FASTA or CSV).

4 / 5

Distinctiveness Conflict Risk

The ToolUniverse-gap framing and format-specific triggers create a clear niche with mostly distinct triggers. Not 5: 'downloading and parsing scientific data' on its own could overlap with generic data-processing or download-helper skills.

4 / 5

Total

18

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
mims-harvard/ToolUniverse
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.