Content
53%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
The body has a solid core — a verified executable script, accurate API examples, and well-organized one-level-deep references — but it is buried in duplicated commands, circular self-references, dated machine-generated paths, and generic boilerplate sections. The workflow is sequenced but its checkpoints are abstract rather than concrete.
Suggestions
Cut the boilerplate sections (Risk Assessment, Security Checklist, Lifecycle Status, Evaluation Criteria, Response Template) and the circular 'See ## X above' cross-references, and deduplicate the py_compile command to a single occurrence.
Fix the command examples: remove the bogus `--help` flag (the script has no argparse and ignores it), document that `python scripts/main.py` runs a canned demo, and delete the dated `cd "20260318/..."` path.
Ground the Workflow steps in concrete validation, e.g. 'run `python -m py_compile scripts/main.py`, then execute the advisor with the confirmed parameters and verify the recommendation's assumptions section before returning it'.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | The ~250-line body is noticeably padded: the py_compile command appears three times (Quick Check, Audit-Ready Commands, Example Usage), the frontmatter description is repeated verbatim in 'When to Use', several sections contain only circular cross-references ("See `## Prerequisites` above for related details"), and boilerplate sections (Risk Assessment, Security Checklist, Lifecycle Status, Evaluation Criteria, Response Template) add no task-specific value. This matches 'Noticeably verbose; several unnecessary explanations or padded sections'; it is not 1 because genuinely useful content (Capabilities, Usage code, Input Parameters, References) is interleaved with the padding. | 2 / 5 |
Actionability | The Usage section shows copy-paste-ready code (StatisticalAdvisor.recommend_test / check_assumptions / calculate_power) that exactly matches the real API in scripts/main.py, plus a concrete parameter table and runnable verification commands — fitting 'Mostly executable guidance; concrete code or commands with minor gaps'. It is not 5 because `python scripts/main.py --help` is misleading (the script has no argparse, so --help is silently ignored and the demo runs instead), the bare `python scripts/main.py` runs a canned demo rather than the documented input-driven run plan, and the example `cd "20260318/scientific-skills/..."` path is a wrong machine-generated artifact. | 4 / 5 |
Workflow Clarity | The Workflow section lists a coherent 5-step sequence with stop-early and fallback rules, matching 'Steps listed but validation gaps; sequence present but checkpoints missing or implicit'. It is not 4 because the checkpoints are abstract policy statements ('validate that the request matches the documented scope') rather than concrete verify commands tied to steps, and the triplicated compile-check sections blur the actual entry path; it is above 2 because a real, ordered sequence with explicit error-handling behavior exists. | 3 / 5 |
Progressive Disclosure | The bundle structure is sound: SKILL.md acts as the overview, detailed material lives in three real, substantive, one-level-deep files (references/statistical_tests_guide.md, assumption_tests.md, power_analysis_guide.md — each clearly signaled under References), and executable logic is separated into scripts/main.py. This fits 'Good structure; most content appropriately placed; references mostly clear; minor organization gaps'. It is not 5 because the body also inlines substantial boilerplate (risk, security, lifecycle, evaluation sections) that adds navigation noise rather than earning its place in SKILL.md. | 4 / 5 |
Total | 13 / 20 Passed |