Content
78%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
A tight, honest instruction-only skill whose publish-time checklist is well sequenced and validated, but it stops short of being executable end-to-end: nothing tells Claude how to actually run the baseline, train the model, or invoke the claim check.
Suggestions
Add the concrete commands (or script paths) for running the mean-pose baseline, training each path, and invoking `ruview_claim_check`, so the workflow is executable rather than advisory.
Briefly define the three split types (chronological / blocked-gap / grouped-bucket) or point to where their construction is specified.
Add an explicit recovery step after the claim check — e.g. what to do when it flags an untagged or 100% claim — to complete the validate-fix-retry loop.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | Lean and efficient — every line carries project-specific, non-obvious discipline (the retraction history, the ~50% mean-pose baseline fact, the leakage-free split taxonomy) and nothing re-explains concepts Claude already knows. | 5 / 5 |
Actionability | The publish checklist is concrete (report delta in pp, run `ruview_claim_check`, tag MEASURED-EQUIVALENT only with the reproducer), but the core train/evaluate workflow has no commands or code, and key details like what `ruview_claim_check` is or how to construct each split are missing — matching the 'some concrete guidance but incomplete' anchor. | 3 / 5 |
Workflow Clarity | The 'Before you publish a number' section is a clear 4-step sequence with an explicit validation checkpoint (step 3 flags untagged or perfect claims), matching anchor 4; it falls short of 5 because there is no error-recovery loop and the training/evaluation steps themselves are not sequenced. | 4 / 5 |
Progressive Disclosure | At 29 lines with no need for external references (no references/, scripts/, or assets/ exist in the bundle), the three well-organized sections satisfy the rubric's under-50-lines exception for a top progressive-disclosure score. | 5 / 5 |
Total | 17 / 20 Passed |