CtrlK
BlogDocsLog inGet started
Tessl Logo

testland/junit-xml-analysis

Parses JUnit-format XML reports (the de-facto interchange format every CI ingests - Jenkins, GitHub Actions, GitLab, Buildkite, CircleCI) into structured, machine-readable per-suite and per-case metrics tables (passed / failed / errored / skipped, time, classname, message, stack), groups failures by classname for trend analysis, and distinguishes "new failures vs flakes" by cross-referencing the `flakyFailure` and `rerunFailure` rerun elements. Use when the downstream consumer is a dashboard, script, or aggregator - not when the goal is a human-readable prose summary (use test-run-summary-author for that). Single-run, in-XML aggregation only; for cross-run cross-environment roll-ups, use a cross-run test-suite aggregator.

90

Quality

90%

Does it follow best practices?

Impact

Average score across 10 eval scenarios

SecuritybySnyk

Passed

No findings from the security scan

Overview
Quality
Evals
Security
Files

SKILL.md

name:
junit-xml-analysis
description:
Parses JUnit-format XML reports (the de-facto interchange format every CI ingests - Jenkins, GitHub Actions, GitLab, Buildkite, CircleCI) into structured, machine-readable per-suite and per-case metrics tables (passed / failed / errored / skipped, time, classname, message, stack), groups failures by classname for trend analysis, and distinguishes "new failures vs flakes" by cross-referencing the `flakyFailure` and `rerunFailure` rerun elements. Use when the downstream consumer is a dashboard, script, or aggregator - not when the goal is a human-readable prose summary (use test-run-summary-author for that). Single-run, in-XML aggregation only; for cross-run cross-environment roll-ups, use a cross-run test-suite aggregator.

junit-xml-analysis

Overview

The "JUnit XML" format is the de-facto schema every CI consumes, emitted by virtually every test runner (pytest, Jest, Vitest, Go test, Maven Surefire, Cypress, Playwright, and the rest).

Per llg-junit (the community schema reference used by Jenkins's parser):

"Root element: <testsuites> (optional if only one suite exists; <testsuite> can be the root instead)."

The hierarchy is testsuitestestsuitetestcase, with result child elements (<failure>, <error>, <skipped>) hanging off each testcase. This skill covers parsing the format, building per-suite + per-case metrics, and the flaky-vs-new distinction via the modern <rerunFailure> / <flakyFailure> extensions.

When to use

  • The CI emits JUnit XML and the team needs PR-time analytics (failure clusters, slow-test list, flake suspects) without buying a commercial test analytics SaaS.
  • A single team has multiple frameworks (pytest in services, Jest in frontend, Go test in tools) - JUnit XML is the lowest common denominator across them.
  • A regression dashboard needs structured input.

Step 1 - Schema overview

Per llg-junit:

LevelRequired attributesCommon attributes
testsuites(none required at root)tests, failures, errors, disabled, time, name
testsuitename, testsfailures, errors, skipped, time, timestamp, hostname, id, package
testcasename, classnametime, assertions, status

Each <testcase> contains at most one of:

  • <skipped message=""> - test not executed
  • <error message="" type=""> - "unanticipated problem (uncaught exception, crash)" (llg-junit)
  • <failure message="" type=""> - "explicit test failure (assertion failed)" (llg-junit)

Plus optional:

  • <system-out> - stdout captured during execution
  • <system-err> - stderr captured during execution
  • <properties> - environment settings as name/value pairs

Critical distinction: per llg-junit, <failure> is an assertion failure (the test made a claim that came back false). <error> is an exception or crash before the assertion ran. Group them differently in dashboards - errors are usually environment / infra; failures are usually code or fixture drift.

Step 2 - Parse safely

Use a streaming parser for large files (multi-thousand-test suites are common). Python core:

# scripts/parse_junit.py
import xml.etree.ElementTree as ET

def parse_junit(path):
    tree = ET.parse(path)
    root = tree.getroot()
    suites = root.findall('testsuite') if root.tag == 'testsuites' else [root]

    for suite in suites:
        for case in suite.findall('testcase'):
            fault = case.find('failure')
            if fault is None:
                fault = case.find('error')
            yield {
                'suite': suite.get('name'),
                'classname': case.get('classname'),
                'name': case.get('name'),
                'time': float(case.get('time') or 0),
                'status': classify(case),
                'failure_message': fault.get('message') if fault is not None else None,
            }

def classify(case):
    if case.find('failure') is not None: return 'failure'
    if case.find('error')   is not None: return 'error'
    if case.find('skipped') is not None: return 'skipped'
    return 'pass'

An Element with no children is falsy, so case.find('failure') or case.find('error') would skip a childless <failure>; test the nodes with is None instead.

Always handle both root shapes: the root may be <testsuites> or a bare <testsuite>. The Node.js (fast-xml-parser) equivalent, which also has to undo single-element collapsing (one testcase = bare object, multiple = array), is in references/junit-xml-parsing.md.

Step 3 - Distinguish new failures from flakes

Per llg-junit, the schema "supports modern variants including <flakyFailure>, <flakyError>, <rerunFailure>, and <rerunError> elements for additional test run metadata."

When the runner does automatic retries (Maven Surefire's rerunFailingTestsCount, pytest-rerunfailures, etc.):

  • A test that passed on retry emits a <flakyFailure> (or <flakyError>) child with the original failure.
  • A test that failed on every retry emits one or more <rerunFailure> children plus the final <failure>.

Classification:

def reliability(case):
    has_flaky = case.find('flakyFailure') is not None or case.find('flakyError') is not None
    has_rerun = case.find('rerunFailure') is not None or case.find('rerunError') is not None
    has_final = case.find('failure') is not None or case.find('error') is not None
    if has_flaky and not has_final:    return 'flaky'        # passed on retry
    if has_rerun and has_final:        return 'consistently_failing'
    if has_final:                      return 'newly_failed'
    return 'pass'

Surface flaky tests in a separate report - they're noise to the PR author but signal to the test-suite owner.

Step 4 - Aggregate per-suite metrics

from collections import defaultdict

def per_suite(cases):
    agg = defaultdict(lambda: {'pass': 0, 'failure': 0, 'error': 0, 'skipped': 0, 'flaky': 0, 'time': 0.0})
    for c in cases:
        agg[c['suite']][c['status']] += 1
        agg[c['suite']]['time'] += c['time']
    return agg

Step 5 - Trend analysis (cross-run)

To detect "is this a new failure or has this test been failing for a week?", store every run's parsed metrics in a per-suite history file:

{"sha":"abc123","ts":"2026-05-05T14:00:00Z","suite":"checkout","failure":2,"flaky":1,"time":12.4}
{"sha":"def456","ts":"2026-05-05T14:30:00Z","suite":"checkout","failure":2,"flaky":0,"time":12.1}

Compare by suite + classname:

classnamenamelast 5 runs resultfirst failed sha
cart.CartTestaddItem_validatesStockF F F F Fabc123 (5 days ago)
checkout.PromoTestapplyPromo_caseInsensitiveP P P P Fthis PR (suspected regression)

The first row is a stale failure; the second is a probable regression.

Step 6 - Per-case slow-test list

Sort testcases by time descending. The top 1% is the fast feedback target - moving any one of them from 30s → 3s saves more than refactoring a hundred tests that already run in <100ms.

Step 7 - CI integration

Run tests with a JUnit reporter enabled, then parse and upload the results on if: always() - JUnit XML matters most on failed runs, so a step gated on success would drop exactly the data you need. The full GitHub Actions workflow (reporter env var, parse step, artifact upload) is in references/junit-xml-parsing.md.

Anti-patterns

Anti-patternWhy it failsFix
Treating <error> and <failure> as the sameErrors are usually infra (DB connection lost), failures are usually code. Conflating hides root-cause patterns.Group them separately.
Dropping <flakyFailure> reports from the dashboardHidden flake budget; quality erodes silently.Surface flaky tests on a separate panel; assign owner.
Loading multi-MB XML with xml.dom.minidom.parseStringWhole-tree-in-memory. OOM on large suites.xml.etree.ElementTree.iterparse for streaming.
Failing the build on any <skipped> count > 0Many runners legitimately skip (platform-gated, conditional).Skip is informational; only fail on failure / error.
Hardcoding <testsuites> as the rootSome runners emit a single <testsuite> as the root.Detect both shapes (Step 2).
Trusting time for sub-millisecond testsSome runners emit 0 for any test under their granularity; sort breaks.Treat time = 0 as "not measured"; don't include in slow-test list.
Cross-suite aggregation by name aloneTwo suites can have a it('renders') each - merging false-flags both.Always group by (classname, name) tuple.

Limitations

  • No standard schema document. "JUnit XML" is a de-facto format with framework-specific dialects. Fields like assertions may or may not appear; the parser must be tolerant.
  • No structured assertion details by default. The <failure> element's body is unstructured text - assertion targets, expected vs actual, and source line are runner-dependent.
  • Time precision varies. Java runners report ms; some Node runners report seconds with 3-decimal precision; some report 0 for fast tests.
  • Reruns require runner support. Frameworks without a built-in retry mechanism don't emit <flakyFailure> - flake detection has to come from cross-run comparison instead (Step 5).

References

  • llg-junit - community-maintained JUnit XML schema reference (used by Jenkins's parser): root element variants, required vs optional attributes, child element catalog including modern <flakyFailure> / <rerunFailure>.
  • coverage-diff-reporter - parallel skill for coverage report diffs (different format, same PR-time analytics shape).
  • allure-reports - richer reporting built on top of allure-results; consumes JUnit XML via per-runner adapters when needed.

SKILL.md

tile.json