CtrlK
BlogDocsLog inGet started
Tessl Logo

testland/mutant-survival-triage

Normalizes a surviving-mutant record across StrykerJS, PIT, mutmut, and Mull into one shape, classifies why it survived (missing case, weak assertion, equivalent mutant, unreachable code, flaky killer), applies per-mutator heuristics for conditional-boundary, arithmetic-operator, statement-removal, and constant mutations, and drafts the specific test that would kill it. Includes the full read-only investigation workflow: take a mutation report (Stryker JSON / PIT XML / mutmut output / Mull JSON) plus the repo at the same commit, read each survivor's mutated line and covering tests, classify, and propose - never auto-rewriting tests. Treats equivalence as a judgment call, because deciding whether a mutant is equivalent to the original is undecidable in general, so a residual survivor rate is expected rather than a defect. Use when a mutation run has finished and the report lists surviving mutants (typically 5+) that nobody has yet explained or turned into concrete test cases.

77

Quality

97%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

Overview
Quality
Evals
Security
Files

Quality

Content

92%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-crafted analysis skill: the missing-case versus weak-assertion decision rule, per-mutator separating inputs, and the equivalent-mutant/undecidability treatment encode genuine domain expertise rather than generic instruction. Structure, progressive disclosure, and workflow sequencing are all strong; the only weakness is length, with a few sections that could be tightened without losing content.

Suggestions

Move the 'If a boundary test already exists and the mutant still survived' discussion out of the Step 4 example template block into its own short subsection, so the template stays a clean copy-paste shape.

Tighten the Step 5 prose: the Stryker/PIT citations and the score-convention guidance could be compressed into a compact list, cutting ~15 lines without losing the load-bearing rules.

The 'Limitations' section partially duplicates content already stated inline (Steps 2, 4, 5); consider merging the coverage-link and operator-coverage caveats into the reference file where the per-tool detail already lives.

DimensionReasoningScore

Conciseness

The body is information-dense (per-mutator separating inputs, tool quirks, score conventions) with little padding, but at ~300 lines it has trimmable sections: the Step 4 example block embeds a discussion paragraph ('If a boundary test already exists...') inside the template, and the Step 5 prose could be tightened. Between anchors 4 and 5 — efficient with minor instances of over-explanation.

4 / 5

Actionability

Fully executable guidance: copy-paste-ready TypeScript for the SurvivedMutant record, four worked mutation families each with concrete separating inputs and proposed fixes, and a complete worked example test (`expect(() => cart.addItem({ qty: 100 })).not.toThrow()`) plus an exact output-format template. This is an instruction-only skill and the guidance is fully actionable.

5 / 5

Workflow Clarity

Clear five-step sequence (normalize → classify → heuristics → propose → equivalence) with an explicit rationale for step ordering, a four-item completeness checklist in Step 4 acting as validation, and error-recovery branching ('If a boundary test already exists and the mutant still survived, either the assertion does not observe the throw... or the production boundary is off by one'). The skill is explicitly read-only ('Never auto-rewrite tests'), so the destructive-operation cap does not apply.

5 / 5

Progressive Disclosure

One reference file (references/tool-normalization.md), clearly signaled with two inline links and a summary of what it carries, exactly one level deep and verified to exist with substantive per-tool tables. The body keeps the workflow and offloads per-tool field mappings and operator-name tables appropriately.

5 / 5

Total

19

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An exemplary description: concrete enumerated actions, tool names and formats as trigger terms, an explicit 'Use when' clause with a specific precondition, and a niche clearly separated from tool-configuration and general test-review skills. Its only cost is length — it is on the far end of the typical description size — but every clause carries either a capability, a trigger, or a boundary.

DimensionReasoningScore

Specificity

Multiple specific concrete actions with comprehensive coverage: 'Normalizes a surviving-mutant record... classifies why it survived (missing case, weak assertion, equivalent mutant, unreachable code, flaky killer), applies per-mutator heuristics for conditional-boundary, arithmetic-operator, statement-removal, and constant mutations, and drafts the specific test that would kill it.' Matches the 5 anchor; the enumerated classes and mutator families leave no minor gaps that would drop it to 4.

5 / 5

Completeness

Explicitly answers both: the what ('normalizes... classifies... drafts the specific test') and the when ('Use when a mutation run has finished and the report lists surviving mutants (typically 5+) that nobody has yet explained or turned into concrete test cases') with a concrete trigger condition.

5 / 5

Trigger Term Quality

Comprehensive natural-term coverage: 'surviving mutants', 'mutation run', 'mutation report', tool names (StrykerJS, PIT, mutmut, Mull), and report formats ('Stryker JSON / PIT XML / mutmut output / Mull JSON'). These are the exact phrases a user with a finished mutation run would say, including tool-name synonyms.

5 / 5

Distinctiveness Conflict Risk

Clear niche (post-run surviving-mutant triage) with distinct triggers — no other plausible skill fires on 'mutation run has finished and the report lists surviving mutants'. Also correctly scopes itself against neighbors ('never auto-rewriting tests', starts at the report, ends at a proposal), further reducing misfire risk. Third-person voice throughout.

5 / 5

Total

20

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Reviewed

Table of Contents