CtrlK
BlogDocsLog inGet started
Tessl Logo

tooluniverse-single-cell

Single-cell RNA-seq analysis with scanpy/anndata — h5ad data loading, scRNA-seq quality control and QC gating (n_genes_by_counts, total_counts, mitochondrial percent / pct_counts_mt, pct_counts_ribo, doublet detection with Scrublet/scDblFinder, ambient RNA / SoupX awareness, empty-droplet filtering, MAD-based thresholds), normalization, dimensionality reduction (PCA, UMAP, t-SNE), clustering (Leiden, Louvain), marker gene identification, cell-type annotation, pseudotime/trajectory analysis. Use for any scRNA-seq workflow, including deciding which cells to filter, flag, or investigate before downstream analysis.

71

Quality

86%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

The canonical home for this skill is tooluniverse-single-cell in mims-harvard/ToolUniverse

SKILL.md
Quality
Evals
Security

Quality

Content

73%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A strong, dense reference skill: executable code, verified helper scripts, an explicit QC-before-analysis sequence with validation gates, and well-organized one-level-deep references. Its main defects are redundancy (duplicated QC guidance with internally inconsistent threshold advice), a broken reference to a missing analysis_patterns.md for the most common DE pattern, and orphaned script files.

Suggestions

Create references/analysis_patterns.md (or fix the three citations pointing to it) so the decision tree's routing for per-cell-type DE and correlation analysis — described as the most common patterns — resolves to real content.

Deduplicate the QC guidance: remove the hardcoded thresholds in 'Interpretation Guidance' (or reframe them explicitly as fallback starting points) so they don't contradict the MAD-based guidance in the QC section, and drop the repeated QC block from 'Complete Pipeline'.

Reference the existing but orphaned scripts (find_markers.py, normalize_data.py, qc_metrics.py) from the relevant body sections, or remove them from the bundle so the file listing matches the documented surface.

DimensionReasoningScore

Conciseness

Mostly efficient — scanpy pipeline code, ToolUniverse tool names, and QC gating rationale are non-obvious content that earns its place — but there is notable redundancy: QC metrics are computed twice (a dedicated QC section plus a 'Complete Pipeline' that repeats simpler QC), and the 'Interpretation Guidance' section restates hardcoded thresholds ("nGenes > 200... < 5000-6000, pct_counts_mt < 20%") that the QC section explicitly warns are only a starting point, creating both padding and mild internal tension. The Scanpy-vs-Seurat table is also largely knowledge Claude already has. This sits at level 3 ('mostly efficient but includes some unnecessary explanation or could be tightened') rather than 4, where over-explanation would be only minor.

3 / 5

Actionability

Guidance is largely copy-paste ready: full scanpy/harmonypy code blocks, exact CLI invocations ("python scripts/scrna_qc.py data.h5ad --doublets", verified against the actual script's flags), and precise ToolUniverse call strings like "tu run run_deseq2_analysis '{...}'". However, the decision tree routes the most common pattern (per-cell-type DE) to "analysis_patterns.md 'Pattern 1'", a file that does not exist in the bundle, so that guidance dead-ends — a concrete gap that keeps it below the level-5 anchor ('specific examples cover the common cases') and at level 4 ('concrete code or commands with minor gaps').

4 / 5

Workflow Clarity

Multi-step processes are clearly sequenced with explicit validation checkpoints: RULE ZERO's pre-computed-results check, a question-to-workflow decision tree, the ordered QC pipeline (empty droplets → per-sample doublets → merge → MAD gating), "Always visualize distributions first" before cutoffs, the per-step removal reporting in scrna_qc.py, the honest-execution stop gate ("do NOT fabricate metrics — print the install plan and stop"), and a troubleshooting table plus synthesis questions for error recovery. This matches the level-5 anchor (clear sequence, explicit validation, feedback loops); level 4 would require missing checkpoints, and the filtering workflow's validations are present.

5 / 5

Progressive Disclosure

Good structure overall: the body is an overview with a decision tree routing to eight clearly-signaled, one-level-deep references (all of which exist) and a Reference Documentation section describing each. The gaps that hold it at level 4 rather than 5: "analysis_patterns.md" is cited three times but missing from the bundle entirely (a broken navigation path), and three of the four scripts (find_markers.py, normalize_data.py, qc_metrics.py) are never referenced from the body, leaving them undiscoverable. It is still well above level 3, where references would be unclear or bulk content inlined.

4 / 5

Total

16

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An exemplary description: concrete and comprehensive capability list, natural trigger vocabulary with synonyms and the h5ad format, an explicit 'Use for...' clause, and a clearly bounded single-cell niche. The dense parenthetical metric list is packed but every term is a concrete capability rather than filler.

DimensionReasoningScore

Specificity

The description enumerates many concrete actions — "h5ad data loading", "QC gating (n_genes_by_counts, total_counts, mitochondrial percent... doublet detection with Scrublet/scDblFinder... MAD-based thresholds)", "normalization", "PCA, UMAP, t-SNE", "Leiden, Louvain" clustering, "marker gene identification, cell-type annotation, pseudotime/trajectory analysis" — with comprehensive coverage of the domain. It matches the level-5 anchor (multiple specific concrete actions, comprehensive) and exceeds level 4, which allows minor coverage gaps; none are evident.

5 / 5

Completeness

It explicitly answers both questions: the 'what' is the detailed capability list, and the 'when' is the explicit clause "Use for any scRNA-seq workflow, including deciding which cells to filter, flag, or investigate before downstream analysis." This is a direct match to the level-5 anchor (clear 'what' AND explicit 'when' with concrete trigger phrases); level 4 would require the 'when' to be less explicit, which it is not.

5 / 5

Trigger Term Quality

Natural user phrasings are well covered with synonyms and the key file format: "Single-cell RNA-seq", "scRNA-seq", "h5ad", "quality control", "filter", "clustering", "marker gene", "cell-type annotation", "PCA, UMAP, t-SNE", "trajectory". This matches the level-5 anchor (comprehensive natural terms including synonyms and file extensions); level 4 requires 'a few natural terms missing', and no meaningful ones are missing.

5 / 5

Distinctiveness Conflict Risk

"Single-cell RNA-seq" with scanpy/anndata, h5ad, and scRNA-specific QC (doublets, ambient RNA, empty droplets) carves a clear niche distinct from bulk RNA-seq or variant-analysis skills, and the trigger terms are scRNA-specific. Minimal conflict risk, matching the level-5 anchor; it is clearly more distinct than the level-4 example, which still carries overlap with a sibling format.

5 / 5

Total

20

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
mims-harvard/ToolUniverse
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.