Content
77%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
An exceptionally actionable, well-sequenced grading workflow whose repo-specific knowledge (taxonomy, surfaces, tags) genuinely earns its tokens. Its weaknesses are repetition of the e2e cost guidance across sections and a monolithic ~200-line body that neither moves hotspot/template detail into reference files nor mentions the bundle's own context-gathering scripts.
Suggestions
Move the Windows/web hotspot lists and the full output-format template into a reference file (e.g., references/report-template.md), keeping only the verdict table and section names inline.
Deduplicate the e2e cost guidance — state it once in Cost guidance and have Constraints reference it in one line instead of restating the rule.
Mention the bundle's scripts (scripts/gather-pr-context.mjs, scripts/gather-local-context.mjs) and how their context.json relates to the "Inputs you'll receive" section so the bundle structure is discoverable.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | Mostly efficient — the test taxonomy, surface tables, and tag mechanics are genuine repo-specific knowledge Claude doesn't have — but the e2e cost guidance is repeated almost verbatim in both "Cost guidance" and "Constraints", and the output template embeds long editorial parentheticals that could be trimmed. It is not a 4 because the duplication and padded sections are noticeable, and not a 2 because most content is non-obvious domain knowledge. | 3 / 5 |
Actionability | Fully executable throughout: copy-paste grep commands (e.g., `grep -r "from.*<filename>" src/vs/**/test/`), exact path patterns per runner, a precise verdict table with fixed emojis, a complete output markdown template, and a two-case decision rule. Every instruction names the concrete file, command, or output line to produce. | 5 / 5 |
Workflow Clarity | The seven investigation steps are explicitly ordered with validation checkpoints (step 4 requires reading candidate tests to confirm they exercise the changed behavior; step 5 conditions on the test surface; step 7 requires grep-confirming a scenario exists before suggesting it), plus a tool-call budget with a defined fallback behavior ("lean toward Insufficient with a note") — matching the anchor for clear sequence with explicit validation and error-recovery handling. | 5 / 5 |
Progressive Disclosure | Sections are well organized and external references (CLAUDE.md, .claude/rules/vitest-tests.md, test/e2e/infra/test-runner/test-tags.ts) are clearly signaled at one level deep, but the body is a ~200-line monolith: the Windows/web hotspot lists and the full output template are inline content that plausibly belongs in reference files, and the three bundle scripts in scripts/ are never mentioned or navigated to from the body. It is not a 4 because the bundle structure itself is undiscoverable and substantial separable content is inlined. | 3 / 5 |
Total | 16 / 20 Passed |