Content
88%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
The body is a strong operational runbook: an explicitly sequenced six-phase investigation with copy-paste-ready az CLI and Python, decision checkpoints, and a defined report format. Its only real weakness is redundancy — the JSON-parsing caveat and artifact-naming details appear both inline and in references/azure-devops-cli.md, costing tokens and slightly blurring the overview/reference split.
Suggestions
Deduplicate the raw_decode JSON-parsing caveat: state it once (in Phase 2 or in references/azure-devops-cli.md, which already documents it) instead of repeating it in the Troubleshooting section.
Move the full artifact-name-to-job mapping list from Phase 3 into references/azure-devops-cli.md (which already lists the naming convention and job names), keeping only the 'TestSummary-simulator_tests<jobname>-1' pattern inline.
Trim the report template's boilerplate headers (e.g., '### Failing Jobs' / '### Infrastructure Issues') to a single compact example block, since the structure is self-evident from the field labels.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | The body is lean: every phase leads with a command or script and prose is limited to genuinely non-obvious operational facts ('The xharness test runner logs are 40K+ lines', 'Each artifact must go to a separate directory'). It is not a 5 because some material is stated twice — e.g., the raw_decode parsing caveat appears in Phase 2 and again in Troubleshooting, and the artifact-name mapping list in Phase 3 restates the naming convention — matching 'efficient; minor instances... that could be trimmed' rather than 'every token earns its place'. | 4 / 5 |
Actionability | Quotes: complete copy-paste-ready 'az pipelines build show', 'az devops invoke ... timeline', and 'az pipelines runs artifact download' commands with a concrete artifact name ('TestSummary-simulator_testsmonotouch_macos-1'), plus fully executable Python for timeline parsing, NUnit XML extraction, and log scanning. Guidance covers the common cases end-to-end, matching the level-5 anchor; level 4's 'minor gaps' does not apply. | 5 / 5 |
Workflow Clarity | Quotes: a six-phase sequence (build overview → timeline → TestSummary artifacts → HtmlReport → raw logs → report) with explicit checkpoints: 'If the build succeeded, tell the user and stop', 'Always start with these before digging into raw logs', 'Only use raw task logs when TestSummary shows BuildFailure', and error-recovery guidance ('If download fails, fall back to log-based analysis'; raw_decode fallback for non-JSON output). This matches the level-5 anchor's clear sequence with validation and feedback loops. The workflow is read-only, so the destructive/batch validation cap does not apply. | 5 / 5 |
Progressive Disclosure | Quotes: 'Read these as needed during investigation' with a well-signaled one-level-deep reference ('references/azure-devops-cli.md ... Read this when you need to construct az commands or download artifacts') — the file exists in the bundle and matches its description. It is not a 5 because the body inlines content duplicated in the reference (artifact naming list, raw_decode caveat, investigation strategy), i.e., 'minor organization gaps' from the level-4 anchor; some of that detail belongs solely in the reference file. | 4 / 5 |
Total | 18 / 20 Passed |