Find flaky tests from CI logs — aggregates pass rates, spots intermittent failures, recommends quarantine. After multiple runs.
64
81%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
Passed
No findings from the security scan
!bash "${CLAUDE_SKILL_DIR}/../../hooks/yaml-helper.sh" resolve_config --keys automation
Automation mode: Resolve modes.automation (project.local.yaml →
project.yaml → default collaborative). Every AskUserQuestion call and
every file write follows .claude/docs/automation-modes.md
(collaborative asks always · guided major-only · autonomous logs and proceeds;
automation_always_ask categories always prompt).
A flaky test is one that sometimes passes and sometimes fails without any code change. Flaky tests are worse than no tests in some ways — they train the team to ignore red CI runs, masking genuine failures. This skill identifies them, explains likely causes, and recommends whether to quarantine or fix each one.
Output: Updated tests/regression-suite.md quarantine section + optional
production/qa/flakiness-report-[date].md
When to run:
/regression-suite identifies quarantined tests that need diagnosisModes:
/test-flakiness [ci-log-path] — analyse a specific CI run log file/test-flakiness scan — scan all available CI logs in .github/ or
standard log output directories/test-flakiness registry — read existing regression-suite.md quarantine
section and provide remediation guidance for already-known flaky testsscan if CI logs are accessible, else
registryCheck for test result artifacts:
ls -t .github/ 2>/dev/null
ls -t test-results/ 2>/dev/nullFor Godot projects: GdUnit4 outputs XML results compatible with JUnit format,
under reports/ (its default report folder, res://reports/). Check reports/,
and any test-results/ you saved runs into, for .xml files.
For Unity projects: game-ci test runner outputs NUnit XML to test-results/
by default.
For Unreal projects: automation logs go to Saved/Logs/. Grep for
Result={Success} and Result={Fail} — each test prints
Test Completed. Result={<status>}
(docs/engine-reference/unreal/current-best-practices.md, "Command Line").
If a path argument is provided, read that file directly.
If no logs found:
"No CI log data found. To detect flaky tests, this skill needs test result history from multiple runs. Options:
- Run the test suite at least 3 times and collect the output logs
- Check CI pipeline output and save a log to
test-results/- Run
/test-flakiness registryto review tests already flagged as flaky intests/regression-suite.md"
Stop and ask the user which option to pursue.
For each CI log or result file found, parse:
JUnit XML format (GdUnit4):
<testcase name= to get test names<failure or <error to identify failuresclassname and name attributes for full test identifiersNUnit XML format (Unity — the file whose <test-run> element /smoke-check
reads):
<test-case element; its fullname attribute is the identifierresult attribute is Passed, Failed, Inconclusive or Skipped
(docs/engine-reference/unity/current-best-practices.md, "Command Line");
only Passed and Failed enter the historyPlain text logs:
PASSED / FAILED adjacent to test namesResult={Success} / Result={Fail}Test passed / Test failedBuild a table: test_id → [run1_result, run2_result, run3_result, ...]
A test is flaky if it appears in the result history with both PASS and FAIL outcomes across runs with no code changes between them.
Flakiness thresholds:
With fewer than 3 runs, every finding is suspected, whatever its fail rate. One failure in two runs reads as 50%, but it is one data point: do not quarantine it, label it suspected, and ask whether more run data is available. The tiers above, and quarantine, apply from 3 runs up.
For each flaky test, classify the likely cause:
| Cause | Symptoms | Fix direction |
|---|---|---|
| Timing / async | Fails after awaiting signals or timers; pass rate correlates with system load | Add explicit await/synchronisation; avoid time-based delays |
| Order dependency | Fails when run after specific other tests; passes in isolation | Add proper setup/teardown; ensure test isolation |
| Random seed | Fails intermittently with no pattern; involves RNG | Pass explicit seed; don't use randf() in tests |
| Resource leak | Fails more often later in a test run | Fix cleanup in teardown; check orphan nodes (Godot) or object disposal (Unity) |
| External state | Fails when a file, scene, or global exists from a prior test | Isolate test from file system; use in-memory mocks |
| Floating point | Fails on comparisons like == 0.5 | Use epsilon comparison (is_equal_approx, Assert.AreApproximately) |
| Scene/prefab load race | Fails when scenes are not yet ready | Await one frame after instantiation; use await get_tree().process_frame |
Use Grep to check the test file for timing calls, randf, global state access, or equality comparisons on floats to narrow down the cause.
For each flaky test:
Quarantine (High flakiness):
"Quarantine this test immediately. Skip it with the engine's own mechanism, log it in the
tests/regression-suite.mdquarantine section, and fix the root cause before removing quarantine."
Skip mechanisms, by engine — name only the one for the project's engine:
func test_x(_do_skip := true, _skip_reason := "flaky: [cause]"). gdUnit4 reads
the argument names do_skip and skip_reason (a leading _ is allowed) in
addons/gdUnit4/src/core/GdUnitTestSuiteScanner.gd, as of gdUnit4 6.1.3;
confirm them there for the installed version.addons/gdUnit4/ source, and docs/engine-reference/godot/ does not
cover them. Log it in the quarantine section and ask the user how their C#
tests are skipped; do not invent an attribute.[Ignore("flaky: [cause]")] — NUnit 3 requires the reasondocs/engine-reference/unreal/ documents no way
to skip an automation test. Log it in the quarantine section and ask the user
how their CI excludes a test; do not invent a flag.Investigate and fix soon (Moderate):
"This test is intermittently unreliable. Root cause appears to be [cause]. Suggested fix: [specific fix based on cause classification]. Do not quarantine yet — fix the test directly."
Monitor (Low/suspected):
"This test shows suspected flakiness. Collect more run data before quarantining. Note it as 'suspected' in the regression suite."
## Flakiness Detection Results
**Runs analysed**: [N]
**Tests tracked**: [N]
### Flaky Tests Found
| Test | System | Fail Rate | Confidence | Likely Cause | Recommendation |
|------|--------|-----------|------------|--------------|----------------|
| [test_name] | [system] | [N]% | confirmed | Timing | Quarantine + fix async |
| [test_name] | [system] | [N]% | confirmed | Float comparison | Fix: use epsilon compare |
| [test_name] | [system] | [N]% | suspected (fewer than 3 runs) | Order dependency | Collect more runs before acting |
### Clean Tests (no flakiness detected)
[N] tests ran across [N] runs with consistent results — no flakiness detected.
### Data Limitations
[Note if fewer than 5 runs were available — fewer runs = less statistical confidence]Ask: "May I update the quarantine section of tests/regression-suite.md
with the flaky tests found?"
If yes: use Edit to append entries to the Quarantined Tests table.
Never remove existing quarantine entries — only add new ones.
Ask (separately): "May I write a full flakiness report to
production/qa/flakiness-report-[date].md?"
The full report includes per-test analysis with cause details and engine-specific fix snippets.
After writing:
is_equal_approx."b21fa0f
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.