CtrlK
BlogDocsLog inGet started
Tessl Logo

python-regression-test-generator

Automatically generates regression tests for Python codebases by analyzing changes between old and new code versions and their existing tests. Migrates tests to work with new code, generates tests for new functionality, and creates mocks for external dependencies. Supports unittest and pytest frameworks. Use when refactoring code, adding features, or ensuring backward compatibility.

64

Quality

76%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/python-regression-test-generator/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

63%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

Highly actionable with concrete, executable examples and a clear seven-step process, but the body is significantly over-long for SKILL.md, duplicating detailed patterns that already live in references/test_generation_patterns.md. Moving the complete and framework-specific examples into the reference (which already covers them) and adding an explicit run-and-verify step would materially improve it.

Suggestions

Trim SKILL.md to the workflow, decision rules (migrate vs create vs obsolete), naming conventions, and constraints; move the 'Complete Example' and 'Framework-Specific Examples' sections into references/test_generation_patterns.md, which already covers much of the same material.

Add an explicit validation step to the workflow (e.g., 'Run the generated tests with pytest/unittest and fix failures before finishing') so quality verification in step 7 is a feedback loop rather than a checklist.

Reduce the ~110 lines of duplicated calculate_total walkthrough to one short example per concept; the async example's mock-setup chain is repeated verbatim three times and can be factored into a helper or shown once.

DimensionReasoningScore

Conciseness

The 560-line body has several padded/duplicated sections: the `calculate_total` example is worked through three times (steps 1–4, the 110-line 'Complete Example', and again in the bundled reference), and the async example repeats the same `aiohttp` mock chain verbatim in three consecutive tests. It also re-explains basics Claude knows (unittest setUp vs pytest fixtures, what to mock). This matches 'Noticeably verbose; several unnecessary explanations or padded sections' rather than the mostly-efficient level-3 anchor.

2 / 5

Actionability

Nearly everything is complete, executable Python: migration examples, mock examples with assertions, setup/teardown in both frameworks, explicit naming conventions, and MUST/MUST NOT constraints. This matches 'Fully executable; copy-paste ready code or commands; specific examples cover the common cases'.

5 / 5

Workflow Clarity

The seven-step workflow (analyze changes → analyze tests → migrate → generate → mock → setup/teardown → quality) is clearly sequenced, and step 7 lists quality checks (deterministic, isolated, executable). It falls short of 5 because there is no explicit feedback loop — no 'run the test suite and re-fix on failure' step with a command — leaving validation implicit rather than a checkpoint.

4 / 5

Progressive Disclosure

A real reference file is listed at the end ('references/test_generation_patterns.md' with its contents summarized), and the body has section structure. However, the large 'Complete Example' and 'Framework-Specific Examples' sections duplicate material in that reference and clearly belong there, fitting 'content that should be separate is inline'. Additionally, references/api_reference.md exists in the bundle but is never referenced from the body.

3 / 5

Total

14

/

20

Passed

Description

88%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: concrete capabilities, explicit 'Use when' triggers, and third-person voice. The main improvement opportunity is broadening trigger synonyms and narrowing the somewhat generic 'refactoring/adding features' triggers.

DimensionReasoningScore

Specificity

The description lists multiple concrete actions — 'analyzing changes between old and new code versions', 'Migrates tests to work with new code', 'generates tests for new functionality', 'creates mocks for external dependencies' — plus framework support ('unittest and pytest'). Coverage of the skill's capabilities is comprehensive, matching the anchor 'Lists multiple specific concrete actions; comprehensive coverage' rather than the level-4 anchor with 'minor gaps'.

5 / 5

Completeness

It explicitly answers what ('Automatically generates regression tests... Migrates tests... creates mocks') and when ('Use when refactoring code, adding features, or ensuring backward compatibility') with concrete trigger phrases, matching the level-5 anchor exactly. It is not level 4 because the 'when' clause is specific and multi-trigger, not merely adequate.

5 / 5

Trigger Term Quality

Natural trigger phrases are present ('refactoring code', 'adding features', 'ensuring backward compatibility', 'regression tests', 'unittest and pytest'), which a user would plausibly say. A few common variations are missing (e.g., 'fix failing tests', 'update tests after changes', 'test migration'), so it fits 'Good keyword coverage; a few natural terms missing' rather than the comprehensive-synonym anchor at 5.

4 / 5

Distinctiveness Conflict Risk

The niche (regression test generation for Python when refactoring) is fairly distinct, but the triggers 'refactoring code' and 'adding features' are broad enough to overlap with general refactoring or testing skills, fitting 'Mostly distinct; minor overlap risk with closely related skills' rather than the minimal-conflict anchor at 5.

4 / 5

Total

18

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (561 lines); consider splitting into references/ and linking

Warning

Total

15

/

16

Passed

Repository
ArabelaTso/Skills-4-SE
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.