Content
48%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
The skill has a clear operational structure (8 well-sequenced operations, explicit state model, real one-level-deep references) but is significantly undermined by verbosity — heavy duplication and generic advice — and by documented script commands that do not match the actual CLI. Roughly 40% of the body could move to the existing reference files without losing anything.
Suggestions
Cut the duplicated 'Quick Reference' section that restates the operations table, task-state tables, best practices, and script commands already covered earlier in the body, and merge the two 'Automation'/'Common Commands' blocks into one.
Fix the documented CLI to match scripts/update-todos.py: replace '--blocker 7 "…"' with '--block 7 --reason "…"', correct '--init task-list.md' to '--init task-list.md my-skill' (two arguments), and remove or correct the nonexistent '--update-estimate' and '--add' flags; also surface the existing '--validate' and '--unblock' flags.
Move the four 'Progress Tracking Formats' sections and the 'Maintain Momentum' strategies into the existing reference guides, keeping SKILL.md as a concise overview with the state model, operations, and script quick-start.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | The ~670-line body is noticeably padded: content is duplicated nearly wholesale (the operations table reappears in 'Quick Reference', the Automation commands reappear as 'Common Commands', best practices appear twice, task-state tables appear twice) and it explains generic todo-management advice Claude already knows ('Take 15-min break', 'Celebrate Progress', 'Remind yourself of MVP'). It is not 1 because the bulk is on-topic skill material rather than extensive background explanation; not 3 because the duplication and generic-advice sections go beyond 'some unnecessary explanation'. | 2 / 5 |
Actionability | The tracking formats and operation processes are concrete with copy-paste markdown examples, but the documented script commands are partially wrong: '--blocker', '--update-estimate', and '--add' do not exist in scripts/update-todos.py (which uses '--block'/'--reason', '--obsolete', '--validate', '--unblock'), and '--init' actually requires two arguments (TASK_FILE SKILL_NAME) while the doc shows one. It is not 4 because a large fraction of the shown commands fail when run; not 2 because the markdown-format guidance and half the CLI are genuinely executable. | 3 / 5 |
Workflow Clarity | Each of the 8 operations has a clear numbered Process with When-to-Use, Output, and examples; checkpoints exist ('Verify Completion: Output exists, quality checked', the Momentum Checklist), and feedback loops are present (blocker → action/escalation, actual-vs-estimated recalibration, in_progress → pending on discovered blockers). It is not 5 because validation is soft in places (e.g., the script's '--validate' flag is never surfaced, and Report Progress has no verification step); not 3 because sequences and checkpoints are explicit throughout. | 4 / 5 |
Progressive Disclosure | The two reference guides (state-management-guide.md, progress-reporting-guide.md) are real, clearly signaled with bold links and one-line descriptions, and one level deep — but the body inlines substantial content that belongs in them (four full tracking-format sections, momentum strategies, extensive duplicated examples, a 200-line Quick Reference). It is not 4 because much content that should be separate is inline despite good reference signaling; not 2 because references are prominent rather than buried and the body has clear section structure. | 3 / 5 |
Total | 12 / 20 Passed |