Content
88%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
The body is highly actionable and the workflow is exemplary — a clearly sequenced six-stage cycle with explicit held-out validation, staging, and backups guarding the destructive adopt step. The main slack is conceptual exposition in the intro and a lack of any references/ split for the flag-table and config detail.
Suggestions
Cut the "It synthesizes three ideas" section and the "deployment-time analogue of training" framing down to one or two lines, or move that background to the skill's README.md.
Move the full CLI flag table and the config-keys list into a references/ file (e.g. references/cli.md), keeping the four core commands (status/dry-run/run/adopt) inline in SKILL.md.
Drop the closing meta-note about the original design-doc path not being vendored; keep only the upstream guide URL, and move the vendoring explanation to README.md.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | The operational core is lean (commands, flag table, config keys), but the conceptual intro ("the deployment-time analogue of training: short-term experience → long-term competence") and the "synthesizes three ideas" section (SkillOpt / Claude Dreams / Agent sleep) explain background Claude does not need to run the skill. This is minor, localized over-explanation — the 4 anchor — rather than the pervasive padding of 3. | 4 / 5 |
Actionability | Commands are copy-paste ready across all operations: "${CLAUDE_PLUGIN_ROOT}/scripts/sleep.sh" status/dry-run/run/adopt, schedule/unschedule, plus a full 14-row flag table with defaults, config keys, and a deterministic validation command (python -m skillopt_sleep.experiments.run_experiment --persona researcher --assert-improves). Matches the 5 anchor for fully executable coverage of common cases. | 5 / 5 |
Workflow Clarity | The six-stage cycle (Harvest -> Mine -> Replay -> Consolidate -> Stage -> Adopt) is clearly sequenced, and validation around the destructive adopt step is explicit: held-out gate, staging where "Nothing live changes", backups on adopt, and "Always show the user the held-out baseline → candidate score ... Evidence before adoption." This matches the 5 anchor with explicit validation and error-recovery guidance (dry-run). | 5 / 5 |
Progressive Disclosure | Sections are well organized and the skill is effectively self-contained (no references/ bundle exists; external detail is one-level-deep via the upstream guide URL). However, the full CLI flag table and config-key reference are inlined where a references/ file would serve better, and the closing meta-note about the non-vendored design-doc path is README material. Fits the 4 anchor: good structure, minor organization gaps. | 4 / 5 |
Total | 18 / 20 Passed |