CtrlK
BlogDocsLog inGet started
Tessl Logo

mutation-testing

Use when evaluating test quality on modules containing business logic, calculation utilities, or state machines to determine whether the tests provide genuine defect detection.

54

Quality

61%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/mutation-testing/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

61%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a well-structured, token-efficient overview with correct progressive disclosure to a single real reference file, and it includes useful concrete heuristics (Stryker, 80% threshold, focused scope). Its weaknesses are the absence of any executable command or code snippet and a Fix workflow without numbered steps or validation checkpoints, plus a mildly padded motivational intro.

Suggestions

Add one minimal executable snippet to the Fix section (e.g., Stryker install command and a scoped run like 'npx stryker run' with a config pointer) so the body is actionable without opening the reference.

Number the Fix steps and add an explicit validation checkpoint (e.g., 're-run the report and confirm the mutation score meets the 80% threshold before finishing').

Trim the opening motivational paragraph to one sentence and connect the references/rule.md pointer to the specific sections that need it (Fix/Code Review).

DimensionReasoningScore

Conciseness

The body is lean overall — a Quick Reference of four dense bullets and four short section stubs — but the opening paragraph over-explains motivation Claude already knows ('A codebase can have 90% line coverage and still ship critical bugs... the standard that actually matters in production'). This fits 'Efficient; minor instances of over-explanation that could be trimmed'. Not 5 because that intro is pure motivational padding; not 3 because the padding is confined to one short paragraph.

4 / 5

Actionability

Some concrete guidance exists — 'Stryker' is named, 'Aim for 80%+ mutation score' and 'Run Stryker on a focused scope (one module)' are specific heuristics — but there are no commands, configuration, or code anywhere in the body ('Set up Stryker for this module, run an initial mutation report' gives no executable steps). This matches 'Some concrete guidance but incomplete; missing key details'. Not 4 because nothing in the body is executable without opening references/rule.md.

3 / 5

Workflow Clarity

The Check → Fix structure supplies a rough sequence ('Set up Stryker... run an initial mutation report... improve tests to kill surviving mutants') and the 80% threshold acts as an implicit checkpoint, but steps are unnumbered, vague ('improve tests' — how?), and there are no validation or feedback-loop steps. This fits 'Steps listed but validation gaps; sequence present but checkpoints missing or implicit'. Not 4 because no explicit validation checkpoint exists.

3 / 5

Progressive Disclosure

The bundle structure is sound: the body is a short overview and implementation detail is correctly split into the real, one-level-deep 'references/rule.md', clearly signaled ('For full implementation details, code examples, and framework-specific guidance, see references/rule.md'). Not 5 because the pointer sits in a generic footer rather than being tied to the relevant sections (e.g., linking Fix and Code Review to the Stryker setup details), and the external rule URL is duplicated in both frontmatter and body.

4 / 5

Total

14

/

20

Passed

Description

62%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description has a strong, explicit 'Use when' clause with a well-scoped trigger context, but it never names its own technique: neither 'mutation testing', 'mutation score', nor 'Stryker' appear, leaving the 'what' vague and the strongest natural trigger terms missing. It is well above the generic bad examples but below the good examples' level of concrete, comprehensive capability listing.

Suggestions

Name the technique and tool explicitly, e.g. 'Run mutation testing with Stryker to measure whether tests catch injected defects...' so the 'what' is concrete.

Add the natural trigger terms users would say — 'mutation testing', 'mutation score', 'Stryker', 'test coverage vs mutation score' — to improve keyword matching.

Consider adding a concrete capability ('set up Stryker, run a mutation report, and improve tests to kill surviving mutants') to raise specificity toward the comprehensive anchor.

DimensionReasoningScore

Specificity

The description names the domain ('evaluating test quality') and one concrete action ('determine whether the tests provide genuine defect detection'), but names no method or tool — 'mutation testing', 'mutation score', or 'Stryker' never appear. It matches the anchor 'Names domain and 1-2 concrete actions, but not comprehensive'. Score 4 would require several specific actions; the single embedded action falls short of that.

3 / 5

Completeness

Both parts are present: an explicit trigger clause ('Use when evaluating test quality on modules containing business logic, calculation utilities, or state machines') and a 'what' ('determine whether the tests provide genuine defect detection'). Not 5 because the 'what' omits the actual technique (mutation testing) and reads more like a purpose than a capability list; not 3 because 'when' is explicit and specific.

4 / 5

Trigger Term Quality

Relevant natural terms are present ('test quality', 'business logic', 'tests', 'defect detection', 'state machines'), but the most common user phrasings — 'mutation testing', 'mutation score', 'Stryker', 'test coverage' — are absent, so a user asking to 'run mutation testing' would not naturally match. This fits 'Some relevant keywords but missing common variations or synonyms'.

3 / 5

Distinctiveness Conflict Risk

The trigger scope (modules with 'business logic, calculation utilities, or state machines') and focus on genuine defect detection make it fairly distinct, with only minor overlap risk against closely related testing skills (e.g., a generic unit-testing or coverage-review skill). It does not meet anchor 5's 'clear niche with distinct triggers' because the absence of the term 'mutation testing' weakens its distinct trigger identity.

4 / 5

Total

14

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

Total

15

/

16

Passed

Repository
thedaviddias/Front-End-Checklist
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.