Content
88%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
A strong orchestrator-style body: fully executable commands for every input shape, a six-phase gated workflow with a threshold-based confidence feedback loop and a Definition of Done checklist, and disciplined deferral of detail to bundle files. The two weaknesses are structural: the phase-to-files mapping is duplicated across two tables (with the confidence policy stated three times), and 8 of the 14 referenced bundle paths (all of `rules/` and `templates/`) are absent from the provided bundle, leaving the core per-phase playbooks unverifiable.
Suggestions
Merge the 'Workflow' phase table and the 'Required reading by phase' table into a single Phase → files → gate table; the two currently duplicate the same rule-file listings per phase and can drift apart.
Ship (or restore) the referenced `rules/*.md` and `templates/analysis-report.md` files in the bundle — 8 of the 14 paths cited in the body resolve to nothing, including the report template required by Phase 5's gate and the Definition of Done.
State the confidence-gate policy once in detail (the 'Confidence-gated iteration' section) and merely cross-reference it from the workflow table and Core principle 5 to remove the triple restatement.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | The body is lean per-sentence — no explanations of concepts Claude already knows (no 'what is Playwright' padding), and details are delegated to bundle files ("Detailed extraction rules, analysis playbooks, and report templates live under `rules/`, `references/`, and `templates/`. Load only what the current phase needs"). However, there is structural redundancy: the "Workflow" phase table's "Rule file" column and the "Required reading by phase" table list the same rule files per phase, and the confidence-gate policy is stated three times (workflow table row 4, the "Confidence-gated iteration" section, and Core principle 5 "Confidence-gated honesty"). This fits the 4 anchor ('Efficient; minor instances of over-explanation that could be trimmed') rather than 5 ('every token earns its place') because merging the two phase tables and de-duplicating the gate policy would remove ~20 lines without information loss; not 3 because the excess is duplicated structure, not unnecessary explanation. | 4 / 5 |
Actionability | Fully executable, copy-paste-ready commands with placeholders and flags: "node <skill_dir>/scripts/trace-extract.mjs <path/to/trace.zip> [--out <dir>]", "node <skill_dir>/scripts/trace-summary.mjs <dir>", "node <skill_dir>/scripts/trace-diff.mjs <pass-dir> <fail-dir>", "node <skill_dir>/scripts/fetch-gh-run.mjs https://github.com/<owner>/<repo>/actions/runs/<id> [--out <dir>]". Concrete detection signals are given ("Magic bytes `50 4b 03 04`", "entries with `type: 'resource-snapshot'`"), plus a model of what a real finding looks like ("'`page.click('text=Save')` waited 4,820ms across 3 attempts'") and a source-mapping mechanism ("action callId → test file/line (Playwright trace events embed `location: { file, line, column }`)"). This matches the 5 anchor; common input cases (zip, dir, JSONL, GitHub Actions URL) are each covered with a specific command, and the `gh`-missing fallback is handled. | 5 / 5 |
Workflow Clarity | A six-phase workflow with an explicit gate per phase and "Do not skip a gate" (e.g. Phase 0 gate: "Format detected, archive unpacked, `trace.trace` + `trace.network` parseable"; Phase 3 gate: "Each hotspot mapped to a code-level cause... with file path or line where possible"), a threshold-based feedback loop ("≥ 90% → Proceed to Phase 5... 70–89% → Run one deeper pass... < 70% → Surface the gap to the user with a question — do not propose changes on speculation"), an iteration cap ("After two deep-dive iterations without reaching 90%, stop and present findings as a hypothesis"), and a "Definition of Done" checklist for final verification. This matches the 5 anchor (clear sequence, explicit validation steps, feedback loops, checklist). The operations are read-only analysis, so the destructive/batch validation cap does not apply. | 5 / 5 |
Progressive Disclosure | The architecture matches the 5 anchor on design: a self-described "thin orchestrator" body with well-signaled, one-level-deep markdown links organized by phase ("Load on demand — do not preload"), and all six files in the provided bundle's `references/` and `scripts/` directories cited by the body exist (flake-patterns.md, performance-patterns.md, trace-extract.mjs, trace-summary.mjs, trace-diff.mjs, fetch-gh-run.mjs). However, scoring against the actual bundle structure per the judging guideline: 8 of the 14 referenced paths — `rules/input-detection.md`, `rules/measurement-methodology.md`, `rules/action-timing.md`, `rules/network-analysis.md`, `rules/console-and-errors.md`, `rules/flake-diagnosis.md`, `rules/confidence-loop.md`, and `templates/analysis-report.md` — are not present in the bundle (no `rules/` or `templates/` directories exist). Since every phase's method and the final report template live in those files, navigation fails for the most load-bearing references, which drops this to the 4 anchor ('good structure; most content appropriately placed; references mostly clear; minor organization gaps') rather than 5 ('easy navigation'). Not 3: the signaling and one-level-deep organization are excellent, and every referenced file that exists in the bundle resolves correctly. | 4 / 5 |
Total | 18 / 20 Passed |