Content
56%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
The skill has a genuinely useful, nearly executable worked example and a logical topic organization, but it is padded with textbook statistics Claude already knows plus generic template boilerplate, and its primary reference file is missing from the bundle. Conciseness and bundle hygiene are the main weaknesses.
Suggestions
Delete the textbook explanations in 'Implementation Details' (test-selection mappings, per-test effect-size table, APA element list) that duplicate both Claude's existing knowledge and the referenced guide files, leaving the SKILL.md as a lean overview pointing to `references/`.
Fix the dangling reference: `references/test_selection_guide.md` is cited as the 'primary decision aid' three times but is absent from the bundle — either add the file or rework the body to not depend on it.
Trim the generic template sections ('Required Inputs', 'Output Contract', 'Input Validation', 'User Checkpoints') to skill-specific guidance, and add `seaborn` to Dependencies since `scripts/assumption_checks.py` imports it.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | The 'Implementation Details' section extensively re-explains textbook statistics Claude already knows ('normal + equal variances → Student's t-test', 'non-normal/ordinal → Mann–Whitney U', 't-tests: Cohen's d', 'ANOVA: partial η²', 'chi-square: Cramér's V'), duplicating content that the referenced files already cover. Generic template sections ('Required Inputs', 'Output Contract', 'Input Validation', 'User Checkpoints') add further padded, skill-agnostic boilerplate, matching the 'noticeably verbose; several unnecessary explanations or padded sections' anchor. | 2 / 5 |
Actionability | The example is concrete and nearly runnable end-to-end (synthetic data, assumption checks, t-test with effect size, power analysis, APA output string), and the script's `comprehensive_assumption_check(data, value_col, group_col, alpha)` signature matches the actual bundle code. Minor gaps prevent a 5: the import is hedged ('If your repo provides this module... otherwise comment it out') and `scripts/assumption_checks.py` imports `seaborn`, which is absent from the Dependencies list. | 4 / 5 |
Workflow Clarity | Implementation Details sections 1–5 mirror the analysis sequence (select test → check assumptions → effect size → power → report), the worked example demonstrates that order, and validation checkpoints exist (assumption check before inference, 'Quick Validation' section, warnings against post-hoc power). It is not 5 because no explicit stepwise workflow is written out for the guided-selection flow itself — the sequence must be inferred from the example and section ordering. | 4 / 5 |
Progressive Disclosure | References are clearly signaled and one level deep, but the body's self-described 'primary decision aid' — `references/test_selection_guide.md` — does not exist in the bundle despite being cited three times, a substantive navigation defect. Additionally, the inline test-selection mappings in 'Implementation Details' duplicate content that belongs in that (missing) reference file, matching the 'some structure but could be better organized' anchor rather than the minor-gaps anchor of 4. | 3 / 5 |
Total | 13 / 20 Passed |