Content
70%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
The body is a dense, well-sequenced playbook with genuinely valuable tool-specific knowledge (parameter names, gotchas, evidence semantics) and a clear 10-phase workflow with fallbacks and a validation checklist. Its weaknesses are inlined textbook-genetics explanations that pad the token budget, non-executable workflow sketches, and a monolithic structure that should push the per-tool API reference into separate reference files.
Suggestions
Move the per-tool API details (parameters, defaults, aliases) for each phase into references/ files (e.g., references/orphanet-tools.md, references/clinvar-tools.md), keeping SKILL.md as a workflow overview with one-line tool summaries and clear links.
Trim the explanations of knowledge Claude already has — the inheritance-mode definitions, the basic consequence hierarchy, and the long deep-intronic parenthetical — down to one-line pointers on how they change the filtering strategy.
Add one fully executable Python example showing how to call a ToolUniverse tool (e.g., Orphanet_search_diseases) from a Bash-run script, so the 'COMPUTE, DON'T DESCRIBE' directive is backed by a copy-paste template.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | Mostly efficient — the bulk is genuinely tool-specific knowledge (parameter gotchas like 'name (NOT query)', association-type semantics, GenCC tiers, review-star confidence levels) — but it includes unnecessary explanation of concepts Claude already knows: the four inheritance-mode definitions ('Autosomal dominant: look for heterozygous variants in ONE copy...'), the basic consequence hierarchy, and a long padded parenthetical on deep intronic variants. This is more than the 'minor instances' of the level-4 anchor, so level 3 fits best. | 3 / 5 |
Actionability | Concrete, executable guidance throughout: exact parameter names, defaults, and aliases for ~25 tools, explicit gotchas, and two example workflows with real call signatures like Orphanet_search_diseases(name="Marfan syndrome"). It falls short of the copy-paste-ready level 5 because the workflows are tool-call sketches, not runnable Python — the 'COMPUTE, DON'T DESCRIBE' section instructs running Python via Bash but never shows how to invoke a ToolUniverse tool from code. | 4 / 5 |
Workflow Clarity | The 10 phases (0-9) are clearly sequenced with an explicit ordering rationale ('phenotype -> disease -> gene -> variant, not the reverse'), fallback strategies provide error-recovery loops when tools return nothing, and the Completeness Checklist and evidence-grading tiers act as end-of-workflow validation. No destructive or batch operations apply a cap. This matches the level-5 anchor; level 4 would lack the checklist or fallback loops. | 5 / 5 |
Progressive Disclosure | Structure is present and well-organized (phase headers, examples, mistakes, limitations), but this is a 283-line monolithic SKILL.md with no bundle files — the per-tool API reference (parameters, defaults, aliases for ~25 tools) is inlined in the body when it clearly belongs in references/ files, matching the level-3 anchor 'content that should be separate is inline' rather than level 4's 'most content is appropriately placed'. | 3 / 5 |
Total | 15 / 20 Passed |