Validate that an Antithesis property catalog can actually catch bugs. For each safety property, inject one realistic mutant into the SUT, run it, and confirm the targeted property fires. Diagnoses survivors as a bad mutant, bad oracle, bad workload, or bad property, and routes fixes back to antithesis-workload or antithesis-research. Use once the harness is built and a baseline run is green.
Skill version: 2026-09-23 a8fe900
Validate the oracle, not the system under test.
A green run shows nothing bad was observed — not that the property would have noticed. Mutation testing removes that ambiguity: for each property, inject one realistic bug designed to break exactly that property, run it under Antithesis, and confirm the property fires.
Success means:
references/evidence-and-report.md, "Verdicts and qualifiers")antithesis-researchantithesis/scratchbook/mutation-testing/report.md records the outcome for every in-scope property, with the run that proves itclean.sh confirms it. Authorized fixes may land there; a mutation never doesUse the antithesis-research skill to build the property catalog, the
antithesis-setup skill to scaffold the harness, and the antithesis-workload
skill to implement assertions and test commands. Use the antithesis-launch
skill to submit runs — do not run snouty launch directly. Use the
antithesis-triage skill to read results.
antithesis/scratchbook/) is missing, or has fallen behind the assertions in the code, reconstruct a catalog from those assertions plus the baseline run — see references/catalog-reconstruction.md. This is not a blocker; it does narrow what the sweep can claim.manifests/ subdirectory. Kubernetes harnesses are not supported yet: mutant selection works by building one image per mutant and swapping the tag in docker-compose.yaml. Tell the user this and stop. Check this before the compose check below — a Kubernetes harness has no docker-compose.yaml.docker-compose.yaml for Antithesis and no manifests/ directory either. Use the antithesis-setup skill to create the compose harness.antithesis-workload skill first — there is no oracle to validate, and a reconstruction cannot stand in for this one.snouty is not installed. See https://raw.githubusercontent.com/antithesishq/snouty/refs/heads/main/README.md for installation options.A green baseline at the current code state is required before any mutant is designed. Mutation testing on a buggy SUT measures nothing: you cannot tell a property that fired because of your mutant from one that was already failing.
Green means: every safety-class property (Always, AlwaysOrUnreachable,
Unreachable) passes.
Separately, check which of the properties you intend to mutate actually appear
in the run's property list. This needs the catalog-slug to property-name
mapping — a run reports the assertion's message string as name, not your
slug — so build it now, per references/sweep-and-verdicts.md, "Map slugs to
property names first". An assertion Antithesis never cataloged cannot be
falsified; this does not fail the gate. Scope each missing one out as
outstanding — not cataloged and continue with the rest
(references/catalog-reconstruction.md diagnoses why it is missing) — except on
a fallback-SDK or JavaScript service, where absence means only "not hit in this
run": mutate those like any other, reading the marker from run events
(references/catalog-reconstruction.md, the join table). Only if
none of them appear is the harness broken rather than incomplete: stop and
report that.
Check the instrumentation signals too — Software was instrumented, Symbols were uploaded, Thread pausing was enabled (antithesis-triage,
references/instrumentation.md). Report them, but gate on them only for a
language that supports coverage instrumentation. Python is cataloging-only and
the fallback SDK emits none; there the fuzzer has no coverage feedback to steer
toward a divergence, so prefer a longer re-run to a bad oracle diagnosis for a
survivor, and say so in the report.
Green does not require every property in the run to pass:
Reachable or Sometimes does not block the sweep: record it, route it to antithesis-workload, and carry it into scoping — a safety property in the same unexercised region is likely to come back outstanding — workload gap.Hypervisor utilization, Customer output volume, and similar) are not SUT correctness signals. Ignore them.A failing safety property stops the skill. Report the failures and route the
user to antithesis-workload or antithesis-research — but read
counterexample_count first. Always and Sometimes imply Reachable
(antithesis-triage, references/properties.md), so a safety property that was
never reached can surface as failing without any invariant having been violated.
counterexample_count above zero is a real violation and stops the sweep; no
counterexamples and no examples means never evaluated — a coverage gap, handled
like the failing Reachable above.
Read status.md for a recorded baseline; it is valid only while its
base_tree fingerprint matches the current one. On a first sweep, create the
file here in the shape references/evidence-and-report.md defines and write to
it continuously — a sweep that only writes it at the end has nothing to resume.
If there is no valid baseline, establish one per
references/sweep-and-verdicts.md, "Establishing the baseline".
When the catalog is being reconstructed, this run comes first and supplies the property list. The gate itself is unchanged — a red safety property still stops the skill, whether or not there was a catalog naming it.
The ceiling is denominated in runs: a measured baseline wall clock far off the interview's rule of thumb changes the schedule, not the authorization — proceed within the ceiling without re-asking.
mNN-<slug> id.references/sweep-and-verdicts.md, "Attributing a falsification").base_tree: The git tree hash of the fork's mutation-base commit. The fingerprint that decides whether a recorded baseline still applies.Reachable assertion where the SDK catalogs assertions; a log record where it does not.Scan first — question 2 is quoted in terms of N, the safety-class property count, and only the scan supplies it — then ask all five before doing anything else. These decide how much the skill spends, how far it goes without you, and what it may change in your tree.
Scan the source for SDK assertion callsites, always — not only when the
catalog is missing (references/catalog-reconstruction.md gives the
per-language spellings). It tells you which of three situations you are in:
references/catalog-reconstruction.md); the scan's safety-class count is the mutant countShow the user the property list either way, with the class of each — the moment to correct a name or scope call is before any budget is spent.
references/mutation-harness.md).references/sweep-and-verdicts.md).antithesis-research, or refine the
catalog in place — correct the property, or withdraw it when it is
conceptually unfalsifiable — and keep going on the others.Answering autonomous plus refine plus apply gives a workflow that runs until it converges or exhausts its budget.
Write the answers to antithesis/scratchbook/mutation-testing/interview.md
before spending anything — the paths, the five answers, and the agreed
scope (references/evidence-and-report.md has the shape), updated whenever an
answer changes. It is the only record of what the user authorized, and what a
stopped sweep resumes from.
The run ceiling confirmed here is the standing authorization for this
session. Later rounds within it need no re-confirmation, in either workflow; a
later session re-confirms once on resume. A round that would exceed the ceiling
stops — to ask, in the checkpointed workflow; to write the report with what is
outstanding, in the autonomous one. Track runs spent against the ceiling in
status.md from the first launch.
Use the antithesis-documentation skill to access these pages. Prefer snouty docs.
https://antithesis.com/docs/concepts/properties_assertions/assertions.mdhttps://antithesis.com/docs/reference/sdk.mdhttps://antithesis.com/docs/product/fault_injection.md| Reference | When to read |
|---|---|
references/mutation-harness.md | Always — the fork, the patch layout, tags, builds, verification |
references/catalog-reconstruction.md | The scratchbook is missing or lags the code — rebuilding a catalog to mutate against |
references/mutant-design.md | Designing a mutant for a property; the three-part patch |
references/static-validation.md | Tracing the kill chain before spending a run |
references/sweep-and-verdicts.md | Launching, polling, classifying survivors, and the iteration loop |
references/evidence-and-report.md | Recording verdicts, and the final report |
references/resume.md | interview.md exists from a previous sweep — resuming it |
interview.mdreferences/mutation-harness.md; copy the harness into antithesis/scratchbook/mutation-testing/, resolve the paths each script takes (--source, --patches, --images, --fork), and materialize the forkantithesis/scratchbook/mutation-testing/images.txt with the compose image names built from this repo (references/mutation-harness.md, "Tags")references/catalog-reconstruction.md and reconstruct one from the assertions and the baseline runreferences/mutant-design.md and references/static-validation.mdreferences/static-validation.md). Review the batch and reject weak, crashy, or un-killable candidates before wiring anymut/<id> branch each, and export them with sync-patches.shverify-mutant.shreferences/sweep-and-verdicts.md; launch the sweep, one select-mutant.sh → launch at a time, with runs overlapping up to the agreed limitreferences/evidence-and-report.md; record verdicts and write the reportclean.sh when the sweep has converged or been abandoned, and after any failure — not when it will be resumed, and a ceiling stop with outstanding properties counts as to-be-resumed: the fork holds any unexported mutant branch, and its base_tree is what a resume diffs against (references/resume.md). If you must clean first, run sync-patches.sh; the resume then treats any fingerprint drift as real and re-baselinesreferences/sweep-and-verdicts.mdsync-patches.sh before re-forking; fork.sh refuses otherwise — never clear that with --force (references/mutation-harness.md, "When a script stops on an assumption")clean.sh when the sweep is finishedBudget exhausted, checkpoint declined, session ended — whatever stopped it,
resume from the recorded state rather than re-deriving it. If interview.md
is absent this is a first sweep: run the opening interview instead. Otherwise
read references/resume.md and follow it: it reconciles status.md against
the platform, salvages un-exported work from the old fork, decides whether the
recorded baseline still applies, and re-enters the two workflows above at the
right step — without redoing any mutant that already has a verdict at the
current base_tree.
| Assertion class | Treatment |
|---|---|
Always, AlwaysOrUnreachable, Unreachable | Mutate. A green safety property is only negative evidence — the vacuous pass this skill exists to rule out |
Sometimes(cond) | Baseline only. A Sometimes fails only when its condition is false in every timeline, so a subtle, prerequisite-gated mutant leaves it green in the timelines that never hit the prerequisite |
Reachable | Baseline only. A green Reachable already proves the workload drove the system there |
"In scope" means every safety-class property in the catalog — the set the report must account for, whether or not each one ends up with a mutant. Two kinds get a verdict without being mutated, and both belong in the report:
Always and AlwaysOrUnreachable, example_count 0 in the baseline means the assertion never evaluated, and a mutant nothing reaches cannot be killed. One 15-minute randomized run is a sample, though, and it under-samples exactly the deep, rare states most worth validating: confirm with a longer baseline before scoping anything out, then route it to antithesis-workload and record it as outstanding — workload gap. The verdict is provisional: it is the same bad workload work order as the ladder's, governed by question 4 and the ceiling like any other real-tree fix. Check this at scoping time.
Unreachable is the exception: example_count 0 is its healthy state. Whether its guarded region is exercised is exactly what the mutant's marker establishes. Take the class from the assertion scan — the catalog's Type field collapses all three safety classes to Safety, so a catalog read alone would scope out every Unreachablereferences/catalog-reconstruction.md — unless the service uses the fallback SDK or JavaScript, where absence is normal and the property is mutated like any other (the join table there). The same verdict covers a mutant that never passes verify-mutant.sh or whose patch has gone STALE: the sweep could not obtain a result, and says soWhile scoping, note any Sometimes whose condition looks trivially satisfiable
— Sometimes(true) in all but name. Report it as a catalog observation with the
Reachable(...) rewrite antithesis-workload's self-review calls for; never
apply it unattended.
references/sweep-and-verdicts.md, "Iterating").--source and --ephemeral so they cannot be mistaken for regressions in the user's real test history.antithesis/scratchbook/mutation-testing/patches/ — one patch per mutant; the patch set is the list of wired mutantsantithesis/scratchbook/mutation-testing/mutants/{id}.md — per-mutant evidence: target, mistake, kill chain, predicted and actual verdictantithesis/scratchbook/mutation-testing/interview.md — the agreed paths, answers, and scope; what a later session resumes fromantithesis/scratchbook/mutation-testing/status.md — sweep state, baseline record, run ids; resumable while the sweep is in progressantithesis/scratchbook/mutation-testing/history.md — append-only log of every completed run: run id, mutant, target property, attempt, diagnosis, action takenantithesis/scratchbook/mutation-testing/report.md — the outcome for every in-scope property## Falsification section appended to each antithesis/scratchbook/properties/{slug}.mdantithesis/scratchbook/property-catalog.md, its evidence files, and existing-assertions.md, when there was no catalog to start fromBefore declaring this skill complete, review the output artifacts against the criteria below — where sub-agents are supported, with a fresh-context reviewer given the path to this skill file.
interview.md records the paths, the five answers, and the agreed scope, and matches what the user actually authorized — including any change made on resumebase_tree before any mutant was designed, and its run id is recorded in status.mdreferences/evidence-and-report.md, "Verdicts and qualifiers" — a reason on everything but falsified, and the base_tree it was obtained atselect-mutant.sh → launch at once; every completed run has a row in history.mdreport.md names every in-scope property with its verdict and, where falsified, the run id that proves it1fd8470
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.