Content
77%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
A well-structured phase-gated workflow with genuinely executable preflight checks, explicit defaults, and strong validation/escalation loops. The two real weaknesses are the broken bundle layout — the rules/ and templates/ files the thin index depends on are absent, undermining both navigation and the flow-emission guidance — and moderate duplication between the deferral lists and the summary decision table.
Suggestions
Ship the referenced bundle files or fix the links: all 5 `rules/*.md` and 6 `templates/*` paths (and `../e2e-testing/templates/spec.md`) point at files absent from the bundle, so the skill's central thin-index promise currently dangles for more than half its references.
Include one minimal inline Maestro YAML snippet (e.g. a launchApp + tapOn id: + assertVisible triple) so flow emission and healing do not depend entirely on the missing `templates/flow.yaml`.
Merge the "Do not reach for this skill" list and the "Decision flow at a glance" table — they restate the same tdd/e2e-testing/WebView deferral rules twice within ~150 lines of each other.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | The body is largely lean — terse decision tables, command blocks, and defaults with no explanation of concepts Claude already knows. Minor trimming opportunities remain: the "Do not reach for this skill" list and the later "Decision flow at a glance" table repeat the same deferral rules (tdd / e2e-testing / WebView) almost verbatim, and asides like "the per-step token cost is already an order of magnitude below screenshot-driven runners" are rationale Claude does not need. This fits the anchor-4 'efficient; minor instances of over-explanation that could be trimmed' rather than anchor 5's every-token-earns-its-place. | 4 / 5 |
Actionability | Phase 0 gives copy-paste-ready read-only checks (`maestro --version`, `jq '.build | has("e2e")' eas.json`, `xcrun simctl list devices available`), Phase 2 prescribes exact defaults (`maestro test .maestro/<flow>.yaml`, `retries: 1` local / `retries: 2` Cloud, `record_screen: false`), and the locator ladder is a concrete ordered procedure. The gap keeping it from anchor 5: the core emit/heal loop has no inline executable example — every flow and template artifact (e.g. `templates/flow.yaml`, `templates/testid-helper.tsx`) is delegated to files that are not present in the bundle, and the heal step itself is never shown as a concrete command or patch procedure. | 4 / 5 |
Workflow Clarity | Phases 0–3 are clearly sequenced with a mandatory preflight gate, explicit halt-and-ask checkpoints before installs or builds, and defined feedback loops: the heal loop is capped at three attempts, then "run `confidence(analysis)`" and "escalate to the user with the Maestro log, the spec, and the proposed locator changes — do not keep healing blindly." Phase 3 adds verification steps (first-run pass or heal convergence, `test-provenance-guard` on imported helpers, EAS workflow wiring check) and a Definition of done checklist, matching the anchor for clear sequence with explicit validation, error-recovery loops, and checklists. | 5 / 5 |
Progressive Disclosure | The design intent is exemplary — "This `SKILL.md` is a thin index... Do not preload everything — load only what the current phase asks for" with one-level-deep, per-phase pointers — but scored against the actual bundle: only `references/` exists (all 5 referenced reference files are real), while the 5 referenced `rules/*.md` and 6 `templates/*` files (plus `../e2e-testing/templates/spec.md`) do not exist in the bundle, so over half the on-demand-load paths dangle. That is more than the 'minor organization gaps' of anchor 4; a reader following the index hits missing files, which is a structural defect the anchor-3 'could be better organized' band captures. | 3 / 5 |
Total | 16 / 20 Passed |