CtrlK
BlogDocsLog inGet started
Tessl Logo

tooluniverse-proteomics-data-retrieval

Find and retrieve proteomics datasets from MassIVE and ProteomeXchange. Search by species, keyword, or accession; retrieve detailed metadata (instruments, publications, species, PTMs studied). Use for locating public proteomics datasets to reanalyze, comparing instrument/protocol coverage across studies, and pre-download dataset evaluation.

70

Quality

85%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

77%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The skill is highly actionable with a well-sequenced workflow and good failure recovery, but it is a single ~290-line monolith that inlines reference material and repeats itself (taxonomy IDs three times, response-format notes twice). Splitting lookup tables into reference files and de-duplicating would raise both conciseness and progressive disclosure.

Suggestions

Consolidate the NCBI taxonomy ID listings (currently in Key Principles, Phase 0, and the 'Common Species Taxonomy IDs' table) into a single table, and remove the duplicate response-format notes between Phase 1 and the Tool Parameter Reference section.

Move the Tool Parameter Reference, species taxonomy table, and Interpretation Framework into a references/ file (e.g., references/tool_reference.md) and link to them from a lean SKILL.md overview, keeping only the workflow and decision logic inline.

Trim the 'Domain Reasoning: Dataset Quality Assessment' paragraph, which overlaps the Interpretation Framework section and re-explains instrument/quantification concepts (TMT compression, DIA vs DDA) that are already covered in the quality-tier table.

DimensionReasoningScore

Conciseness

The body is mostly dense and useful (parameter tables, decision logic, fallbacks), but there is real duplication and over-explanation: NCBI taxonomy IDs appear three times (Key Principles, Phase 0, and a dedicated 'Common Species Taxonomy IDs' table), response-format notes are repeated in Phase 1 and again in 'Tool Parameter Reference', and the 'Domain Reasoning' paragraph re-explains TMT/DIA/instrument-resolution concepts that largely reappear in the 'Interpretation Framework' section. This fits 'mostly efficient but includes some unnecessary explanation or could be tightened' (3) rather than the minor-trimming of level 4.

3 / 5

Actionability

Every phase gives copy-paste-ready tool calls with exact parameter names, types, and formats (e.g., 'MassIVE_search_datasets(page_size=20, species="9606")', 'species as string, max 100', 'px_id in PXD format only'), plus a complete report template and per-failure fallbacks. Per the scoring notes, an instruction-only skill is not penalized for lacking code; this meets the 'fully executable, specific examples cover the common cases' anchor, not the 'minor gaps' of level 4.

5 / 5

Workflow Clarity

The four phases are clearly sequenced with decision logic (Phase 0 routes accession vs. species vs. keyword queries to the right tools), and the 'Fallback Strategies' table provides explicit error-recovery loops (e.g., 'MassIVE_get_dataset fails for PXD accession -> Use ProteomeXchange_get_dataset instead'), matching the anchor with feedback loops for error recovery. Level 4 would require these recovery paths to be only mostly present; here every failure mode in the flow has a stated fallback.

5 / 5

Progressive Disclosure

There is no bundle at all (no references/, scripts/, or assets/ directories) and the file is ~290 lines: the Tool Parameter Reference, the species taxonomy table, and the Interpretation Framework are reference material inlined in SKILL.md that would fit separate files. Section headers are clear, but content that should be separate is inline with no references signaled, matching anchor 3 rather than the 'content appropriately split' of level 4 or 5.

3 / 5

Total

16

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: concrete and comprehensive on capabilities, explicit about when to use it, and clearly distinguishable from sibling proteomics skills. The only gap is a handful of natural trigger synonyms (mass spectrometry, MS data, PXD/MSV accession formats).

DimensionReasoningScore

Specificity

The description lists multiple concrete actions with enumerated outputs: 'Search by species, keyword, or accession; retrieve detailed metadata (instruments, publications, species, PTMs studied)'. This matches the anchor for multiple specific concrete actions with comprehensive coverage; it goes beyond the 1-2 actions of level 3 and beyond the 'minor gaps' of level 4 since search modes, both repositories, and returned metadata fields are all named.

5 / 5

Completeness

Both questions are explicitly answered: 'what' is covered by the two opening sentences naming the repositories, search modes, and metadata returned, and 'when' is covered by the explicit 'Use for locating public proteomics datasets to reanalyze, comparing instrument/protocol coverage across studies, and pre-download dataset evaluation' clause. This matches the anchor for clearly answering both what and when with concrete trigger phrases, not the weaker 'when' of level 4.

5 / 5

Trigger Term Quality

Good keyword coverage with natural terms users would say: 'proteomics datasets', 'MassIVE', 'ProteomeXchange', 'species', 'accession', 'PTMs', 'public proteomics datasets'. It misses a few common synonyms, notably 'mass spectrometry', 'MS data', and accession formats like 'PXD'/'MSV', so it falls between the 'good coverage, a few natural terms missing' (4) anchor and the comprehensive synonym/extension coverage of 5.

4 / 5

Distinctiveness Conflict Risk

The niche is unambiguous: retrieval of public proteomics datasets from two named repositories (MassIVE, ProteomeXchange). Repository names and accession terminology make it unlikely to fire for analysis-only or other-omics skills, matching the 'clear niche with distinct triggers; minimal conflict risk' anchor rather than the 'minor overlap risk with closely related skills' of level 4.

5 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
mims-harvard/ToolUniverse
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.