CtrlK
BlogDocsLog inGet started
Tessl Logo

tooluniverse

ToolUniverse plugin router. STEP 1 BEFORE ANY ANALYSIS: if the data folder contains `*_executed.ipynb`, run `tu run read_executed_notebook '{"data_folder":"<path>","search":"<keyword>"}'` to extract its cell outputs and apply EVERY filter/sample-exclusion the notebook used — even when the question says 'Using DESeq2/Run X/Compute Y' (this describes the METHOD the notebook used, not a request to rerun). The notebook's cell outputs are the only published authoritative answers; reimplementing or reading stale pre-computed CSVs in the data folder produces different numbers because of outlier-sample removal, library version, and filter steps you don't see by skimming. STEP 2 routing — pick a sub-skill name from this exact list (never invent): tooluniverse-rnaseq-deseq2 (RNA/miRNA-seq DE, correlation, PCA, clustering, dispersion), tooluniverse-gene-enrichment (GO/KEGG/Reactome/GSEA/pathway enrichment), tooluniverse-statistical-modeling (regression, ANOVA, ordinal/logistic, chi-square...

56

Quality

62%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./plugins/tooluniverse/skills/tooluniverse/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

71%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a highly actionable router: exact commands, filename heuristics, a decision-ordered RULE ZERO, and explicit tie-breaking and fallback rules leave no ambiguity about what to do first. Its weaknesses are redundancy — duplicated keyword rows routing the same terms to different skills, and procedural content repeated from the frontmatter — and a monolithic inline structure that keeps everything in SKILL.md rather than splitting the routing table and script catalog into reference files.

Suggestions

Deduplicate the routing table: the peptide-target row appears verbatim in §2 and §4, and 'microbiome'/'16S' route to both tooluniverse-metagenomics-analysis and tooluniverse-microbiome-research with no tie-break — pick one destination or add a rule.

Move the ~25-row Bundled Scripts table and the per-domain keyword tables into reference files (e.g., references/routing-table.md, references/scripts.md), keeping only the routing procedure and the highest-frequency signals in SKILL.md.

Trim the duplicated RULE ZERO prose: the executed-notebook instructions already appear almost verbatim in the frontmatter description; state them once in the body in their full form.

DimensionReasoningScore

Conciseness

The body is mostly operational with almost no explanation of concepts Claude already knows, but it carries clear fat: the peptide-target row appears verbatim in both §2 and §4, "microbiome"/"16S" route to two different skills in §9 (metagenomics-analysis vs microbiome-research), the ~25-row "Bundled Scripts" table cross-references other skills' directories, and RULE ZERO's notebook instructions duplicate the frontmatter description nearly word-for-word. This matches anchor 3 ('mostly efficient but includes some unnecessary content or could be tightened'), below anchor 4 because of the duplicated rows and cross-skill catalog.

3 / 5

Actionability

Guidance is fully executable throughout: exact commands like `tu run read_executed_notebook '{"data_folder":"/path/to/data","search":"<keyword>"}'` and `cd /path/to/data && python3 run_*.py`, copy-paste one-liners for each recurring tool (e.g., `tu run clinical_trial_ae_severity_test '{"dm_file":"DM.csv",...}'`), concrete filename heuristics ("`*_counts.csv`+`*meta*.csv` → RNA-seq DE"), and explicit routing calls `Skill(skill="<name>")`. This matches anchor 5's copy-paste-ready commands covering the common cases.

5 / 5

Workflow Clarity

The routing workflow is a clear numbered sequence (read question AND file list → find keyword row → call Skill → follow loaded skill), RULE ZERO defines an ordered decision procedure (executed notebook → canonical script → own code) with canonical-vs-scratch script disambiguation, and tie-breaking rules plus fallbacks ("If no signal matches → general strategies", "Still ambiguous: Ask user") handle edge cases. It sits at anchor 4 rather than 5 because there are no post-route verification checkpoints and the duplicate keyword rows (two microbiome routes, IC50 in both §4 and §7) create unresolved ambiguities the tie-break rules don't explicitly cover.

4 / 5

Progressive Disclosure

The one bundle reference is properly signaled and one level deep — "load [references/general-strategies.md](references/general-strategies.md)" points to a file that exists — but ~250 lines of keyword routing table plus a cross-skill script catalog are inlined in SKILL.md when they could be split into reference files, matching anchor 3 ('content that should be separate is inline'). Not anchor 4: the routing table and Bundled Scripts table are the bulk of the document and are not separated; not anchor 2: the structure that exists is well-sectioned and the single reference is clearly signaled, not buried.

3 / 5

Total

15

/

20

Passed

Description

53%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description communicates a distinct niche and some concrete commands, but it front-loads ~2,600 characters of procedural instructions (STEP 1/STEP 2) into what should be a capability summary, is written in imperative rather than third-person voice, omits the when-to-use guidance from the description field itself, and is literally truncated mid-sentence ("...ordinal/logistic, chi-square..."). A tighter third-person summary with explicit trigger phrases and the complete sub-skill scope would score substantially higher.

Suggestions

Rewrite the description in third person as a capability summary (e.g., "Routes scientific data-analysis questions to specialized ToolUniverse sub-skills and extracts authoritative results from executed notebooks"), moving the STEP 1/STEP 2 procedural detail into the body where it already exists.

Add an explicit 'Use when...' clause with natural trigger phrases and file extensions to the description field itself, instead of relying on the separate when_to_use field that description-only consumers never see.

Fix the truncated sub-skill list — either complete the enumeration or summarize the scope ("60+ sub-skills covering genomics, clinical decision support, phylogenetics, chemistry, and statistics") so the description does not end mid-sentence.

DimensionReasoningScore

Specificity

The description names several concrete actions — "run `tu run read_executed_notebook ...` to extract its cell outputs" and "pick a sub-skill name from this exact list (never invent): tooluniverse-rnaseq-deseq2 (RNA/miRNA-seq DE, correlation, PCA, clustering, dispersion), tooluniverse-gene-enrichment (GO/KEGG/Reactome/GSEA...)" — which sits at anchor 4, but it is written in imperative second-person voice ("run", "pick") rather than third person, incurring the rubric's −1 voice penalty, and the sub-skill list is cut off mid-sentence at "chi-square...", leaving coverage incomplete. Not anchor 2: it is far more concrete than 'names the domain but actions are minimal'; not anchor 5: the truncation and procedural framing leave real gaps.

3 / 5

Completeness

The 'what' is stated clearly ("ToolUniverse plugin router" plus the notebook-extraction and routing behaviors), but the description field contains no 'Use when...' clause or equivalent trigger guidance — that content is in a separate non-standard `when_to_use` frontmatter field — so per the judging guideline completeness is capped at 3. It is not anchor 4 because the 'when' is absent from the description, not merely implicit; not anchor 2 because the 'what' is concrete, not vague.

3 / 5

Trigger Term Quality

Domain keywords are present and relevant ("differential expression", "pathway enrichment", "regression, ANOVA, ordinal/logistic, chi-square", "PCA, clustering, dispersion"), matching anchor 3's 'some relevant keywords but missing common variations or synonyms'. The natural user-facing trigger phrases ("differential expression", "pathway enrichment", "clinical data", file extensions) live in the separate `when_to_use`/`paths` fields rather than the description itself, and the keyword list is truncated, so coverage falls short of anchor 4.

3 / 5

Distinctiveness Conflict Risk

"ToolUniverse plugin router" with namespaced sub-skills ("tooluniverse-rnaseq-deseq2", "tooluniverse-gene-enrichment", "tooluniverse-statistical-modeling") forms a clear, distinct niche with minimal conflict risk against other skills, matching anchor 4. It falls short of anchor 5 because the router implicitly claims the whole scientific-analysis domain and the description never finishes enumerating its sub-skills, so its boundaries are not fully drawn.

4 / 5

Total

13

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
mims-harvard/ToolUniverse
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.