Content
85%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
An operationally excellent skill: executable code with correct parameter names, response formats, decision thresholds, explicit fallbacks for missing data, and a well-sequenced six-phase workflow with evidence grading. Its main structural weakness is that it is a single monolithic file — the per-tool API details belong in a reference file with SKILL.md kept as the overview and workflow guide.
Suggestions
Move the per-tool parameter documentation and response-format details (Phases 1-5 code blocks, ~150 lines) into a references/tool-api.md file, keeping SKILL.md as the phase overview with one or two key calls per phase and a clearly signaled link (e.g., "See [references/tool-api.md](references/tool-api.md) for all tool parameters and response formats").
Deduplicate the parameter pitfalls: state each gotcha once (either inline in the code comment or in "Common Mistakes and Fallbacks", not both) — the MONDO-vs-EFO rule and several param-name warnings currently appear two to three times.
Extract the tool-by-tool notes (e.g., "GTEx uses v8 data (v10 endpoints may return empty)", "gwas_get_associations_for_trait is BROKEN") into a compact reference table so the workflow sections carry only phase-level reasoning.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | Code blocks are dense with tool-specific gotchas, return shapes, and thresholds that Claude would not know, so most tokens earn their place; however, pitfalls are stated redundantly (MONDO-not-EFO appears in section 2e, Phase 5, and Common Mistakes; param-name warnings appear both inline and in "Common Mistakes"), which is a minor trim opportunity. Not 5 because of that repetition; not 3 because the bulk is non-obvious operational knowledge rather than padding. | 4 / 5 |
Actionability | Fully executable ToolUniverse calls with correct parameter names ("variant_id, NOT rsid"), expected response formats, decision thresholds (CADD >=20, L2G > 0.5, RegulomeDB 1a-2a), and a working REST fallback with pandas aggregation — copy-paste ready and covering the common cases. Matches the top anchor. | 5 / 5 |
Workflow Clarity | Six phases are explicitly sequenced with an overview map, and error recovery is built in: "Response format variable: list, {data, metadata}, or {error} -- handle all three", a "When data is missing, adapt the analysis" fallback table, broken-tool warnings, and evidence-graded synthesis with confidence tiers. The workflow is read-only analysis, so the destructive/batch validation cap does not apply; the fallback and error-handling guidance constitutes the feedback loops the top anchor requires. | 5 / 5 |
Progressive Disclosure | No bundle files exist (references/, scripts/, assets/ are absent) and the entire 370-line skill is monolithic: roughly 150 lines of per-tool parameter documentation (Phases 1-5 code blocks) are reference material inlined in SKILL.md, and no external reference is signaled (the lone "See tooluniverse-data-wrangling skill" mention points to a sibling skill, not a bundle file). Section structure is clear, which keeps it above 2, but "content that should be separate is inline" fits anchor 3 better than 4, where references would be mostly present and clear. | 3 / 5 |
Total | 17 / 20 Passed |