CtrlK
BlogDocsLog inGet started
Tessl Logo

mutation-testing

Evaluate Python test suite quality using mutmut to introduce code mutations and verify tests catch them. Use for mutation testing, test quality assessment, mutant detection, and test effectiveness analysis.

86

1.17x
Quality

83%

Does it follow best practices?

Impact

87%

1.17x

Average score across 3 eval scenarios

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

82%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A strong, well-structured body: fully executable commands and examples, a clear run-inspect-improve workflow, and disciplined brevity with no padding. The main gaps are minor — an implicit re-run verification step in the improvement loop and a moderately long single file where a reference split could help.

DimensionReasoningScore

Conciseness

The body is efficient — no padded library comparisons or explanations of basic concepts; the Key Concepts section defines mutmut-specific vocabulary (Killed/Survived/Mutation Score) rather than general programming knowledge, and code examples are purposeful. It misses level 5 because the opening paragraph partially re-explains what Key Concepts already covers, and the ~50-line weak-vs-strong example block could be trimmed, matching 'efficient; minor instances of over-explanation that could be trimmed'.

4 / 5

Actionability

Quick Start gives copy-paste-ready commands (mutmut run/results/show/apply), the Configuration block is a complete ini example, the Python examples are executable and cover the common case, and the CI yaml is directly usable — matching 'fully executable; copy-paste ready code or commands; specific examples cover the common cases'. Nothing is pseudocode or abstract.

5 / 5

Workflow Clarity

A clear sequence exists (run → results → show → apply → reset) with the results-inspection acting as a validation checkpoint, and the Improving Mutation Score section provides a fix loop for survived mutants. It stops short of level 5 because the loop is not explicitly closed — it never says to re-run mutmut to confirm the new tests kill the mutant — matching 'clear sequence with most checkpoints present; minor validation gaps'. Validation steps are present, so the missing-validation cap does not apply.

4 / 5

Progressive Disclosure

No bundle files exist, and the body is well-organized with clear section headers (Key Concepts, Quick Start, Configuration, operators table, CI, Best Practices) and no nested or dead references, matching 'good structure; most content is appropriately placed; minor organization gaps'. It does not reach level 5 because at ~170 lines some content (the CI workflow and mutation operator table) arguably belongs in a separate reference file, but it clearly exceeds level 3 since nothing is buried or clearly misplaced.

4 / 5

Total

17

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: it names the tool and concrete actions, uses third-person imperative voice, and pairs a clear 'what' with an explicit 'Use for...' trigger clause containing multiple natural trigger phrases. The only weaknesses are minor — slight overlap risk from 'test quality assessment' and a few missing natural synonyms like 'kill mutants'.

DimensionReasoningScore

Specificity

Phrases like "Evaluate Python test suite quality using mutmut", "introduce code mutations", and "verify tests catch them" name the domain, the tool, and several concrete actions (evaluate, introduce, verify), matching the 'several specific actions; minor gaps' anchor. It is not level 5 because coverage is not comprehensive (no mention of reviewing results, improving tests, or CI integration), and not level 3 because it goes beyond 1-2 actions by naming the tool plus three distinct operations.

4 / 5

Completeness

It explicitly answers both questions: the "what" ("Evaluate Python test suite quality using mutmut to introduce code mutations and verify tests catch them") and the "when" ("Use for mutation testing, test quality assessment, mutant detection, and test effectiveness analysis") with concrete trigger phrases. This matches the level-5 anchor exactly (parallel to the PDF exemplar), so neither 4 (weaker 'when') nor any lower anchor applies.

5 / 5

Trigger Term Quality

The trigger clause "Use for mutation testing, test quality assessment, mutant detection, and test effectiveness analysis" contains the natural phrases users would say, matching the 'good keyword coverage; a few natural terms missing' anchor. It falls short of 5 because common variations like "kill mutants", "coverage gaps", or tool-adjacent phrasings are absent, but it is clearly above level 3's 'missing common variations' since four distinct natural terms are present.

4 / 5

Distinctiveness Conflict Risk

Mutation testing with mutmut is a clear niche with distinct triggers, but the phrase "test quality assessment" could overlap with generic test-writing or code-review skills, matching 'mostly distinct; minor overlap risk with closely related skills'. It does not fit level 5 because that overlap risk is more than minimal, and level 3 would understate how tool-specific the description is.

4 / 5

Total

17

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
mixpanel/mixpanel-headless
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.