Content
63%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
A highly actionable, well-structured operational guide with concrete commands and clear monitoring/recovery flows, weakened by token-wasting repetition, jokey padding, inline time-sensitive trivia, and a monolithic single-file layout where per-agent detail could be split into reference files.
Suggestions
Consolidate the PTY guidance into one prominent statement and drop the jokey asides ('read your soul docs', 'Skippy gets pinged') and the Learnings section, which restates rules already covered — this alone would remove roughly 30–40 lines of the conciseness penalty.
Move time-sensitive facts (the gpt-5.2-codex default model, Pi's PR #584 prompt-caching note dated Jan 2026) into a clearly labeled version/status section so they can be pruned without touching the core workflow.
Split the per-agent sections (Codex flags, Claude Code, OpenCode, Pi) into references/ files (e.g. references/agents.md) and keep SKILL.md as the pattern plus tool-parameter overview, giving the bundle one level of clearly signaled references.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | The body is mostly tables and code, but there is unnecessary padding: the PTY warning is repeated at least four times ('Always use tty:true', rules #1, the Learnings section, plus inline comments), and jokey asides ('it'll read your soul docs and get weird ideas about the org chart', 'Skippy gets pinged in seconds, not 10 minutes') cost tokens. Time-sensitive detail ('Pi now has Anthropic prompt caching enabled (PR #584, merged Jan 2026)', 'gpt-5.2-codex is the default') is inline rather than in a dated/deprecated section, which the guidelines penalize. The Learnings section largely restates rules already given above. This fits anchor 3 — mostly efficient but could be tightened — and not anchor 4, whose 'minor instances' understates the repetition. | 3 / 5 |
Actionability | Guidance is largely copy-paste ready: 'exec_command tty:true workdir:~/project background:true command:"codex exec --full-auto ..."', 'write_stdin session_id:XXX chars:""', parameter tables, and concrete JSON subagent task payloads. Minor gaps keep it below anchor 5: line 113 runs 'codex --yolo ...' without 'exec' (which would open interactive mode rather than one-shot), and 'codex review --base origin/main' is presented without confirming it is an actual subcommand. Anchor 4 — 'concrete code or commands with minor gaps' — is the best fit. | 4 / 5 |
Workflow Clarity | Sequences are clear: start in background → poll with write_stdin → send input → kill_session, and the PR review flow (worktree isolation preferred, manual clone fallback, cleanup) is well ordered with error-recovery rules ('If an agent fails/hangs, respawn it or ask the user for direction'). Verification is delegated to the subagent prompts ('run focused checks', 'report changed files plus verification') rather than an explicit orchestrator-side checkpoint confirming agent output before reporting done — a minor validation gap matching anchor 4, not the explicit validate/re-validate loops of anchor 5. | 4 / 5 |
Progressive Disclosure | The skill has no bundle files at all (no references/, scripts/, or assets/), and ~270 lines of per-agent detail (Codex flags, Pi/OpenCode/Claude Code sections, Learnings) live inline in SKILL.md. Sections and headers are well organized, but content that would naturally split into per-agent reference files is inlined — matching anchor 3 ('some structure... content that should be separate is inline') rather than anchor 4, since there are no one-level-deep references to point to. | 3 / 5 |
Total | 14 / 20 Passed |