Content
73%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
A well-structured, disciplined instruction skill with an explicit workflow, validation gates, a QA checklist, and genuine one-level-deep references. The main costs are redundant restatements of audit-mode and blocker rules, and dangling references to a non-existent examples/ directory.
Suggestions
Consolidate the three read-only-audit-mode passages (intro paragraph, 'Non-negotiable quality bar' exception, and the 'Read-only audit mode' section) into one authoritative section and reference it elsewhere in one line.
Merge the overlapping blocker content between step 1 validation, the failure-mode policy, and the final QA gate so each constraint (subject-x-task repeated measures, missing seeds, contradictory interpretation) is stated once.
Either add the examples/ files listed under 'Example files' or remove that section; also consider inlining or relocating the ../research-ideation/references/research-contract.md cross-skill reference so all paths resolve inside the skill bundle.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | No padding with concepts Claude already knows (no explanation of p-values, CIs, or plotting libraries), but the same rules are stated repeatedly: read-only audit mode is specified in the intro, again in a dedicated section, and echoed in the quality-bar exception; the subject-x-task repeated-measure blocker appears in step 1 and again in the failure-mode policy; quarantine rules appear twice. This is more than the minor trimming the 4 anchor describes, though never vague or padded enough to fall to 2. | 3 / 5 |
Actionability | Concrete, instruction-level guidance throughout: exact output filenames and directory tree, an explicit per-figure requirement list (purpose, plotted variables, error-bar meaning, caption requirements, interpretation checklist), a copy-paste-ready claim-candidate markdown template, and specific statistical deliverables (mean +/- std, 95% CI, effect sizes, multiple-comparison handling). Falls short of 5 because the statistical execution itself is delegated to references and some directives remain abstract ("use non-parametric fallback when assumptions fail" without naming the fallback tests). | 4 / 5 |
Workflow Clarity | A clearly sequenced 6-step workflow with explicit validation checkpoints: step 1 validates artifacts and unit of analysis with a hard stop ("If the comparison is not statistically valid, say so before continuing"), step 2 locks comparison questions before running statistics, and the Final QA gate is an explicit do-not-finish-until checklist. Feedback loops for error recovery are present via the failure-mode policy and the quarantine rule for contradictory statistics files, matching the 5 anchor rather than the 4 anchor with minor validation gaps. | 5 / 5 |
Progressive Disclosure | The body is an overview with six real, on-topic reference files (all verified present in references/) listed with one-line purposes under "Load only what is needed" and signaled in-context at the relevant workflow steps — solid one-level-deep structure. However, the "Example files" section lists three files (examples/example-analysis-report.md, example-stats-appendix.md, example-figure-catalog.md) that do not exist in the bundle, and one reference path points outside the skill directory (../research-ideation/references/research-contract.md), which is unverifiable here — keeping it below the 5 anchor but above the 3 anchor, since the defect is confined to peripheral paths. | 4 / 5 |
Total | 16 / 20 Passed |