Decides what kind of failure a red CI test is before anyone starts fixing it. Extracts seven failure signals from the runner output, stack trace, run history, and environment metadata, then walks an ordered first-match-wins rule set to exactly one verdict: flaky-known, environment-drift, defect, timeout, flaky-pre-incident, or flake-of-unknown-cause. Emits the verdict together with the alternatives that were rejected and the specific condition each one failed, so the triage decision is auditable rather than asserted. Use when a test has just gone red and the next action depends on whether the cause is a product defect, a non-deterministic test, or drifted infrastructure.
98
89%
Does it follow best practices?
Impact
99%
1.12xAverage score across 10 eval scenarios
Passed
No findings from the security scan
Triage document produced
100%
100%
Pricing failure separated from the outage group
80%
96%
Pricing failure identified as a real regression with its cause
83%
87%
MUST NOT classify from the commit subject alone
100%
100%
MUST NOT issue a single verdict for the whole build
100%
100%
Outage group correctly grouped and dispositioned
90%
100%
Rerun plan answered directly
100%
100%
Rejected explanations recorded with observed values
100%
100%
Next action names the owner and preserves the reproduction
77%
100%
Triage document produced
100%
100%
Service image change identified as the cause
100%
100%
Application changes excluded by the replayed commit
100%
100%
MUST NOT recommend rotating, re-saving, or changing the secret
100%
100%
MUST NOT file this as a defect against an application team
100%
100%
Rejected explanations recorded with observed values
100%
100%
Next action is pinning and provisioning, routed to the image owner
100%
100%
One call, not a hedge
100%
100%
Triage document produced
100%
100%
States that a two-run-old test cannot be settled from this material
100%
100%
MUST NOT choose between the wrong-selector and missing-feature stories
100%
100%
The two runs identified as non-comparable
100%
100%
Missing artefacts named specifically
100%
100%
Correctly excludes what can be excluded
100%
100%
Next action is evidence collection, and the agent does not invent its result
100%
100%
No routing to a team the evidence does not support
100%
100%
Triage document produced
100%
100%
Termination identified as memory exhaustion, not a timeout
100%
100%
Cause attributed to the runner-size change
100%
100%
Checkout commit explicitly excluded using run 9829
100%
87%
MUST NOT route this to the checkout team as a product defect
100%
100%
Rejected explanations recorded with observed values
100%
100%
Next action is the runner configuration, owned by CI
100%
100%
One call, not a hedge
100%
100%
Triage document produced
100%
100%
States plainly that the material does not settle the cause
100%
100%
MUST NOT assert the dependency drift as the cause
100%
100%
Missing artefacts named specifically and usefully
100%
100%
Distinguishes unavailable evidence from evidence against
100%
100%
Establishes what the log does show
100%
100%
Next action collects evidence rather than applying a fix
100%
100%
No verdict laundered through confidence language
100%
100%
Triage document produced
100%
100%
Identified as shared-infrastructure contention, not a payments problem
100%
100%
Evidence drawn from beyond the failing test
100%
100%
Time clustering established from the history
100%
100%
MUST NOT route this to the payments team as a defect
100%
100%
MUST NOT prescribe a retry, a quarantine, or a longer timeout
100%
100%
Rejected explanations recorded with observed values
100%
100%
Next action addresses the pool or the schedule
100%
100%
Triage document produced
100%
100%
Identified as a product regression despite the timeout shape
100%
100%
The commit named as the cause with the mechanism
100%
100%
MUST NOT treat the wait-shaped failure as evidence of flakiness
100%
100%
Prior failures examined rather than counted
100%
100%
The wiki page rejected as a formal record
60%
100%
Direct answer on the retry wrapper
100%
100%
Rejected explanations recorded with observed values
100%
100%
Next action preserves the reproduction
50%
100%
Triage document produced
0%
100%
Identified as an intermittent coupling inside the suite
0%
100%
Worker co-location pattern extracted from the history
0%
100%
MUST NOT classify this as an environment or infrastructure failure
0%
100%
MUST NOT offer a rerun as the disposition
0%
100%
Rejected explanations recorded with observed values
0%
100%
Worker-count change assessed rather than blamed
0%
100%
Next action is confirmation then the team's intermittent-test process
0%
100%
Triage document produced
100%
100%
Cause identified as the merged change, not the environment
100%
100%
Timeline used to refute the image roll
100%
100%
Pinned-image rerun used as evidence
100%
100%
MUST NOT recommend pinning the runner image
100%
100%
Rejected explanations recorded with observed values
100%
83%
Routed to the owner of the change
100%
83%
Direct recommendation given on the open PR
100%
100%
Triage document produced
100%
100%
Invoice failure identified as a real regression
100%
100%
MUST NOT treat the source comment or issue #4412 as evidence of flakiness
100%
100%
Notifications failure identified as already-quarantined
100%
100%
Different handling for the two, stated as such
100%
100%
Rejected explanations recorded with observed values
100%
100%
Next actions differ and neither is a bare rerun of the invoice test
100%
100%
No stacked verdicts
100%
100%
Table of Contents