Content
88%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
A strong orchestrator body: routing and loop workflows are gated by executable scripts with explicit exit-code semantics, and every rule is checkable. The main costs are mild redundancy between the Hard rules and Anti-patterns sections, an inlined forcing-question library that could live in a reference, and referenced bundle paths that are not present in the skill directory to verify.
Suggestions
Deduplicate guidance repeated across sections — the single-participant-anecdote rule and the DORMANT-escalation rule each appear in both Hard rules and Anti-patterns; state each once and cross-reference.
Move the six-entry forcing-question library (with its canon citations) into a reference file (e.g. references/forcing_questions.md) and keep one lane-defining example per lane in SKILL.md to reduce inline token cost.
Ship the referenced bundle files (scripts/product_goal_router.py, scripts/discovery_cadence_tracker.py, scripts/ost_linter.py, the three references/*.md, assets/sample_discovery_log.json) alongside SKILL.md so the one-level-deep reference structure actually resolves.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | The body is dense and mostly assumes competence (exit-code tables, command blocks, terse rules), but there is some redundancy that could be trimmed: "no insight is asserted from a single participant (anecdote, not insight)" reappears in Anti-patterns ("promote a single-participant anecdote to insight"), as does "DORMANT escalates by name" / "escalate to the product lead by name", and the six-entry forcing-question library with full canon citations is verbose for inline placement. This sits between level 4 (efficient, minor instances that could be trimmed) and level 5 (every token earns its place) — noticeably above the midpoint but not flawless, so 4. | 4 / 5 |
Actionability | Fully executable, copy-paste-ready commands with arguments and interpreted exit codes: `python3 scripts/product_goal_router.py --text "<the goal>" --output json` (exit 0/2/3 behaviors spelled out), `python3 scripts/discovery_cadence_tracker.py --input discovery_log.json` (exit 5 refusal), `python3 scripts/ost_linter.py --input ost.json` (exit 2 = NEEDS-REWORK), and the full goal_compiler invocation with --manifest and --out. This matches level 5 (fully executable commands covering the common cases); level 4 would require missing key details, and even log shapes are surfaced via "every tool ships --sample" and the assets sample path. | 5 / 5 |
Workflow Clarity | Both multi-step processes have clear sequences with machine-checkable validation gates and feedback loops: routing branches on exit codes (0 → load skill, 2 → ask one clarifying question, 3 → restate, "Never guess silently"), and the discovery loop is five numbered steps with lint-before-cite enforcement ("exit 2 = NEEDS-REWORK, fix before citing the tree"), explicit stop states, and named escalation. This matches level 5 (clear sequence, explicit validation steps, feedback loops, checklists). | 5 / 5 |
Progressive Disclosure | Structure is good: three clearly signaled one-level-deep references in a References section (continuous_discovery_canon.md, product_operating_model.md, ai_product_evals.md) plus inline pointers to scripts and an assets sample, and the When-to-invoke table delegates detail to the sub-skills. However, no bundle files (references/, scripts/, assets/) are present in the skill directory, so the referenced paths cannot be verified, and substantial content (the forcing-question library with citations, the full hard-rules detail) is inlined where an orchestrator overview could push it one level down — matching level 4 (good structure, references mostly clear, minor organization gaps) rather than level 5. | 4 / 5 |
Total | 18 / 20 Passed |