Content
63%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
The body is a well-templated, largely actionable EDD guide, but it is held back by duplication (grader types, pass@k guidance, and storage layout each appear twice), placeholder steps in the evaluate workflow, and progressive-disclosure problems: the ~300-line monolith inlines detail that belongs in referenced files, and its file references point to bundle paths that are absent. Consolidating the duplicates and either shipping or removing the dangling references would lift it substantially.
Suggestions
Deduplicate the repeated material: merge 'Product Evals (v1.8)' grader types and pass@k guidance into the earlier 'Grader Types' and 'Metrics' sections, and collapse 'Eval Storage' with 'Minimal Eval Artifact Layout'.
Add an explicit failure-handling loop to the Evaluate phase (e.g. 'if any eval fails: fix the regression, re-run `/eval check <feature>`, only report when all pass'), and replace placeholder steps like '[Run each capability eval, record PASS/FAIL]' with concrete commands.
Fix the reference topology: either ship the referenced files (scripts/eval-harness.js, scripts/lib/eval-harness/, docs/architecture/eval-harness-frameworks.md) or remove the pointers, and move the dense 'Local Framework Utilities' prose into a one-level-deep reference file to slim SKILL.md.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | Mostly efficient — templates, commands, and formats rather than concept explanations — but several sections are padded or duplicated: 'When to Activate' restates the frontmatter triggers, the grader-type list appears twice ('Grader Types' and again under 'Product Evals (v1.8)' as 'Code grader / Rule grader / Model grader / Human grader'), pass@k guidance appears twice ('Metrics' and 'pass@k Guidance'), and the eval storage layout is described twice ('Eval Storage' and 'Minimal Eval Artifact Layout'). This fits the level-3 anchor ('mostly efficient but includes some unnecessary explanation or could be tightened'); it is not level 4 because the duplication is substantive rather than minor, but not level 2 since nothing explains concepts Claude wouldn't know. | 3 / 5 |
Actionability | The body provides mostly executable guidance: copy-paste grader commands ('grep -q "export function handleAuth" src/auth.ts && echo "PASS" || echo "FAIL"'), CLI invocations ('node scripts/eval-harness.js example', '/eval check feature-name'), and complete markdown eval templates. It is not level 5 because some steps remain placeholders, e.g. '[Run each capability eval, record PASS/FAIL]' and '[Write code]' in the workflow example, leaving minor gaps for a fresh user. | 4 / 5 |
Workflow Clarity | The four-phase sequence (Define → Implement → Evaluate → Report) is clearly laid out with a worked 'Example: Adding Authentication', and regression evals act as a built-in checkpoint. It is not level 5 because the Evaluate phase lacks an explicit failure feedback loop — there is no 'if an eval fails, fix and re-run' step, and the workflow example defers to '/eval check' rather than spelling out validation; it is well above level 3 since sequence and most checkpoints are present. | 4 / 5 |
Progressive Disclosure | The body is sectioned with clear headers, but the bundle contains no references/, scripts/, or assets/ directories, so the inline pointers ('scripts/lib/eval-harness/', 'node scripts/eval-harness.js example', 'See docs/architecture/eval-harness-frameworks.md') reference files that do not exist in the skill, and substantial material (framework utility details, the duplicated Product Evals v1.8 content) is inlined where it could live one level deep. This matches the level-3 anchor ('some structure but could be better organized; references present but not clearly signaled; content that should be separate is inline') — structure exists, but the reference topology is broken and content placement needs work. | 3 / 5 |
Total | 14 / 20 Passed |