Content
92%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
A high-quality operational skill: fully executable commands, an explicitly ordered workflow with validation checkpoints at every phase, error-recovery feedback loops, and a clean one-level reference split (parsing rules, manifest schema, error matrix) with all referenced files present and non-nested. The only notable weakness is redundant restatement of platform/HybridApp/exit-code boundaries across multiple sections, which trims a little token efficiency without adding information.
Suggestions
Consolidate the boundary declarations: state the mWeb/Desktop `EXIT 4` and HybridApp `EXIT 7` gates once (e.g., in the exit-codes table) and reference them from Scope and 'Out of scope' rather than restating the full rationale in each place.
Remove the duplication between the 'Exit codes' table and references/error-handling.md — keep the table in the body for exit-vs-continue semantics and move the per-situation triggers fully into the reference, or vice versa.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | The body is dense and operational — no tutorial prose or explanations of concepts Claude already knows, and every section carries task-specific information. However, boundaries are repeated across sections: the mWeb/Desktop decline with `EXIT 4` appears in Scope, triage gates, and 'Out of scope'; the HybridApp gate is stated in the intro, Scope, and 'Out of scope'; the exit-code semantics are duplicated between the exit-codes table and references/error-handling.md. This fits anchor 4 (efficient with minor instances that could be trimmed) better than anchor 3, since the redundancy is reinforcement of constraints rather than unnecessary explanation. | 4 / 5 |
Actionability | Nearly every operation is given as a copy-paste-ready command: `gh pr view <num> --json title,body`, `agent-device open "$APP_ID" --device "$DEVICE_NAME"`, `agent-device record start "$RUN_DIR/$PLATFORM/flow-$ID.mp4" --fps 24`, `test -s "$TEST_FLOW.ad" || {...}`, exact cache paths, and the fingerprint formula `sha256(precondition + json(steps) + platform)`. Concrete tables (inputs, exit codes, cost guards) cover the common cases; the few prose steps (LLM-driven flow execution) are inherently non-scriptable and are still given precise semantics (verbatim step text, append only successful actions, record chosen values in `params:`). Matches anchor 5. | 5 / 5 |
Workflow Clarity | The sequence is explicit and ordered: numbered triage gates 'run in order, before any device work' → steps parsing → cache check → shared setup → Phase 1 → Phase 2 → manifest → handoff, with per-flow status semantics. Validation checkpoints are present throughout (Phase 1 step 6 sanity-checks the script with `test -s`, step 4 verifies final state via `agent-device is exists`, Phase 2 step 7 verifies the MP4 with `test -s` + `file`), and there are feedback loops and recovery paths (retry Phase 2 once on a 0-byte recording per references/error-handling.md, mark-and-continue per flow, exit codes 5/6 for total failures). This satisfies the batch-operation requirement — the destructive/batch cap of 3 does not apply because validation is explicit — and matches anchor 5. | 5 / 5 |
Progressive Disclosure | The body is an orchestration overview with well-signaled, one-level-deep references, all of which exist: [`references/steps-parsing.md`], [`references/manifest-schema.md`], and [`references/error-handling.md`], each referenced from its matching section ('Steps parsing', 'Manifest schema', 'Error handling') and each self-contained with no nested references. Detail material (parsing heuristics, schema field semantics, the error matrix) is correctly split out while orchestration stays in the main file. This matches anchor 5 (clear overview, well-signaled one-level references, appropriate split). | 5 / 5 |
Total | 19 / 20 Passed |