Content
73%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
A strong, dense reference skill: executable code, verified helper scripts, an explicit QC-before-analysis sequence with validation gates, and well-organized one-level-deep references. Its main defects are redundancy (duplicated QC guidance with internally inconsistent threshold advice), a broken reference to a missing analysis_patterns.md for the most common DE pattern, and orphaned script files.
Suggestions
Create references/analysis_patterns.md (or fix the three citations pointing to it) so the decision tree's routing for per-cell-type DE and correlation analysis — described as the most common patterns — resolves to real content.
Deduplicate the QC guidance: remove the hardcoded thresholds in 'Interpretation Guidance' (or reframe them explicitly as fallback starting points) so they don't contradict the MAD-based guidance in the QC section, and drop the repeated QC block from 'Complete Pipeline'.
Reference the existing but orphaned scripts (find_markers.py, normalize_data.py, qc_metrics.py) from the relevant body sections, or remove them from the bundle so the file listing matches the documented surface.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | Mostly efficient — scanpy pipeline code, ToolUniverse tool names, and QC gating rationale are non-obvious content that earns its place — but there is notable redundancy: QC metrics are computed twice (a dedicated QC section plus a 'Complete Pipeline' that repeats simpler QC), and the 'Interpretation Guidance' section restates hardcoded thresholds ("nGenes > 200... < 5000-6000, pct_counts_mt < 20%") that the QC section explicitly warns are only a starting point, creating both padding and mild internal tension. The Scanpy-vs-Seurat table is also largely knowledge Claude already has. This sits at level 3 ('mostly efficient but includes some unnecessary explanation or could be tightened') rather than 4, where over-explanation would be only minor. | 3 / 5 |
Actionability | Guidance is largely copy-paste ready: full scanpy/harmonypy code blocks, exact CLI invocations ("python scripts/scrna_qc.py data.h5ad --doublets", verified against the actual script's flags), and precise ToolUniverse call strings like "tu run run_deseq2_analysis '{...}'". However, the decision tree routes the most common pattern (per-cell-type DE) to "analysis_patterns.md 'Pattern 1'", a file that does not exist in the bundle, so that guidance dead-ends — a concrete gap that keeps it below the level-5 anchor ('specific examples cover the common cases') and at level 4 ('concrete code or commands with minor gaps'). | 4 / 5 |
Workflow Clarity | Multi-step processes are clearly sequenced with explicit validation checkpoints: RULE ZERO's pre-computed-results check, a question-to-workflow decision tree, the ordered QC pipeline (empty droplets → per-sample doublets → merge → MAD gating), "Always visualize distributions first" before cutoffs, the per-step removal reporting in scrna_qc.py, the honest-execution stop gate ("do NOT fabricate metrics — print the install plan and stop"), and a troubleshooting table plus synthesis questions for error recovery. This matches the level-5 anchor (clear sequence, explicit validation, feedback loops); level 4 would require missing checkpoints, and the filtering workflow's validations are present. | 5 / 5 |
Progressive Disclosure | Good structure overall: the body is an overview with a decision tree routing to eight clearly-signaled, one-level-deep references (all of which exist) and a Reference Documentation section describing each. The gaps that hold it at level 4 rather than 5: "analysis_patterns.md" is cited three times but missing from the bundle entirely (a broken navigation path), and three of the four scripts (find_markers.py, normalize_data.py, qc_metrics.py) are never referenced from the body, leaving them undiscoverable. It is still well above level 3, where references would be unclear or bulk content inlined. | 4 / 5 |
Total | 16 / 20 Passed |