CtrlK
BlogDocsLog inGet started
Tessl Logo

testland/flaky-test-quarantine

Builds a quarantine workflow for flaky tests - marks the test with the framework's skip/fixme/retry annotation, records the failure-rate observation and a bisect link in the annotation body, sets an auto-expiry date, and produces a CI report listing every quarantined test that has expired and needs re-evaluation. Use when a flaky test is blocking the trunk and must be removed from the gating path without losing track of it.

88

1.45x
Quality

84%

Does it follow best practices?

Impact

89%

1.45x

Average score across 10 eval scenarios

SecuritybySnyk

Low

Low-risk findings worth noting

Overview
Quality
Evals
Security
Files

Evaluation results

82%

31%

"Just set retries to five and move on"

Criteria
Baseline
With context

Deliverables produced

100%

100%

Suite-wide retries: 5 refused with a reason

100%

80%

Local retries left at zero

100%

100%

Retry-rescued tests converted into tracked entries

0%

100%

Every record carries a deadline and an owner

0%

100%

Consistently failing test excluded from the retry conversation

0%

0%

One-off deploy-window failure left alone

100%

100%

Never-failing test untouched

100%

100%

Retry budget reconciled against the time budget

75%

100%

100%

40%

A PR that deletes four tests to make the board green

Criteria
Baseline
With context

Deliverables produced

100%

100%

Refund test kept in the codebase

66%

100%

Obsolete test deletion upheld

100%

100%

Remaining two kept but taken out of the blocking path

50%

100%

Every retained-but-inactive test has a dated deadline

23%

100%

Every retained-but-inactive test has a named owner

50%

100%

Record states the measured rate and what happens at the deadline

50%

100%

Untouched test left untouched

100%

100%

96%

30%

Auditing what our integration suite is actually not running

Criteria
Baseline
With context

Deliverables produced

100%

100%

Every remaining exclusion carries a deadline in the annotation itself

16%

100%

Every remaining exclusion names an owner

18%

100%

Re-enabling the test whose blocker closed

100%

100%

Dead commented-out block deleted

100%

100%

Measured rate recorded for the two that stay off

71%

100%

Untracked exclusion given a ticket

100%

66%

Executing test untouched

100%

100%

86%

3%

Five switched-off tests and nobody attached to any of them

Criteria
Baseline
With context

Deliverables produced

100%

100%

Dissolved team never assigned

100%

100%

Unroutable entries escalated to a named decision-maker

87%

100%

No owner invented for the uncovered path

100%

100%

Routable entries resolved correctly

100%

100%

Named individual, not only a team handle, for the entries already past review

50%

50%

Consequence of an unowned entry stated

21%

28%

Existing fields preserved

100%

100%

94%

37%

Nobody has looked at the skip list since March

Criteria
Baseline
With context

Deliverables produced

100%

100%

Q-01 not extended a third time

0%

100%

Blanket ninety-day extension refused

100%

100%

Q-03 returned to the blocking path

50%

50%

Q-04 deleted as dead code

100%

100%

Q-05 extended once, and marked as the last extension

37%

100%

Q-02 left untouched

70%

100%

Extension count carried in the rewritten record

42%

100%

Every surviving entry has an owner and a date

100%

100%

100%

19%

PR author wants his own new test switched off so his PR can land

Criteria
Baseline
With context

Deliverables produced

100%

100%

New test refused

100%

100%

Failure evidence read as a product defect

100%

100%

Pre-existing unrelated flake acted on

100%

100%

Deadline and owner on the entry

0%

100%

Merge condition stated

100%

100%

Passing sibling untouched

100%

100%

Unrelated test untouched

100%

100%

98%

58%

Six red tests and a release branch cut tomorrow

Criteria
Baseline
With context

Both deliverables produced

100%

100%

admin bulk export treated as a regression, not silenced

0%

100%

Sub-threshold test left alone

100%

100%

72% test refused as a flake

0%

100%

Qualifying tests stay in the file and keep running as non-blocking

66%

100%

Every entry carries a re-evaluation deadline

26%

100%

Every entry carries a named owner

50%

100%

Entry records the measured rate, run count, and investigation state

68%

87%

Rejected candidates carry a next action and an owner

50%

100%

65%

-11%

The list of "flaky" tests came out of standup, not out of the data

Criteria
Baseline
With context

Deliverables produced

100%

100%

Zero-failure test refused

100%

100%

Four-run sample refused as a rate

100%

0%

The 0.8% test left alone

100%

100%

Unmentioned worst offender surfaced

100%

100%

Consistently failing test not treated as intermittent

0%

0%

Every recorded rate carries its run count

100%

100%

Entries taken out carry a deadline and an owner

0%

50%

Standup note never used as the basis for an action

100%

100%

100%

25%

Two payment tests failing intermittently, both look the same from the dashboard

Criteria
Baseline
With context

Decision document produced

100%

100%

Capture test refused

100%

100%

Cross-reference to the incident log made explicit

100%

100%

Capture test routed to the right owner as a product defect

100%

100%

Genuine flake correctly taken out of the blocking path

50%

100%

The record for the flake carries a deadline and an owner

0%

100%

Fix identified rather than left open-ended

90%

100%

Third test untouched

100%

100%

73%

43%

We switch tests off and then nothing ever tells us to look again

The agent wrote 45 files outside the evaluated workspace (With context)

These are not visible to scoring or included in the download.

/home/agent/go/pkg/mod/cache/download/gopkg.in/check.v1/@v/list

/home/agent/go/pkg/mod/cache/download/gopkg.in/check.v1/@v/v0.0.0-20161208181325-20d25e280405.mod

/home/agent/go/pkg/mod/cache/download/gopkg.in/yaml.v3/@v/list

/home/agent/go/pkg/mod/cache/download/gopkg.in/yaml.v3/@v/v3.0.1.info

/home/agent/go/pkg/mod/cache/download/gopkg.in/yaml.v3/@v/v3.0.1.lock

Criteria
Baseline
With context

Scheduled mechanism produced

0%

100%

Malformed annotations rewritten into a readable form

0%

90%

Already-stale entries surface on the first run

0%

0%

Nothing is acted on automatically

100%

100%

Report routes to the named owner, not a channel

0%

50%

The e2e job is left alone

100%

100%

Format document is short and shows the fields

0%

91%

Tests left in their current state

100%

100%

Evaluated
Agent
Claude Code
Model
Claude Sonnet 4.6