Content
60%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
The core skill is genuinely good — a clear conditional workflow from PMID lookup to text-analysis fallback to scale selection, backed by real, correctly split bundle files. It is dragged down by heavy auto-generated boilerplate (~100 of ~145 lines add nothing Claude does not already know) and several template artifacts: a truncated helper section, a reference to a nonexistent CONFIG block, a wrong-direction cross-reference, and a machine-specific example path.
Suggestions
Delete the generic template sections (Key Features, Implementation Details, Required Inputs, Output Contract, Validation and Safety Rules, Failure Handling, and the redundant When to Use/When Not to Use) and keep only the Workflow, Helper Scripts, and Quick Validation sections — this alone would move conciseness toward the top anchors.
Complete the truncated PDF Text Extraction section with the actual usage command from extract_pdf.py's docstring (`python scripts/extract_pdf.py <input PDF file> [--output <output file>]`), and remove the reference to the nonexistent in-file CONFIG block in the Example run plan.
Fix the misleading cross-reference 'See ## Workflow above for related details' (the Workflow section is below it) and replace the machine-specific example path (20260316/scientific-skills/...) with a path relative to the skill root.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | Roughly two-thirds of the body is template boilerplate that assumes no intelligence and pads the token budget: "Use this skill when the request matches its documented task boundary", "Packaged executable path(s): scripts/extract_pdf.py plus 1 additional script(s)", "Output discipline: keep results reproducible... avoid undocumented side effects", plus duplicated When to Use / When Not to Use / Required Inputs / Output Contract / Validation and Safety Rules / Failure Handling sections. This clearly matches the 'noticeably verbose; several unnecessary... padded sections' anchor; it is short of level 1 only because it does not explain basic domain concepts (what a PDF or an RCT is), it just restates generic policy. | 2 / 5 |
Actionability | The core workflow is executable: `python scripts/selector.py "<PMID>"` is a copy-paste-ready command, the fallback gives concrete keywords ("Randomized controlled trial", "RCT", ...), the output is a literal JSON block, and the referenced bundle files (references/scale_rules.md, both scripts) all exist. It stops short of level 5 because the PDF Text Extraction section is truncated mid-sentence ("use extract_pdf.py to extract the text content before assessment:" with nothing following) and the Example run plan tells the user to edit an in-file CONFIG block that does not exist in either script. | 4 / 5 |
Workflow Clarity | The four-step workflow is clearly sequenced with an explicit conditional checkpoint (if selector.py returns non-empty JSON, skip to step 3; otherwise fall back to text analysis), which is a genuine validation/branching step. It is not level 5 because there is no error-recovery loop for the PubMed failure path beyond falling back, the 'See ## Workflow above' cross-reference in Implementation Details points the wrong direction (Workflow appears below it), and the truncated helper section and phantom CONFIG reference add confusion. | 4 / 5 |
Progressive Disclosure | Scored against the actual bundle: SKILL.md is an overview, the scale-selection rules are properly split into a real, well-signaled one-level-deep reference ([scale_rules.md](references/scale_rules.md)), and the executable logic lives in the two real scripts. This fits the 'good structure; most content is appropriately placed; references mostly clear; minor organization gaps' anchor rather than level 5, because the truncated PDF-extraction section and the ten generic boilerplate sections dilute navigation and bury the real workflow. | 4 / 5 |
Total | 14 / 20 Passed |