CtrlK
BlogDocsLog inGet started
Tessl Logo

bug-reproduction-test-generator

Automatically generates executable tests that reproduce reported bugs from issue reports and code repositories. Use when users need to: (1) Create a test that reproduces a bug described in an issue report, (2) Generate failing tests from bug descriptions, stack traces, or error messages, (3) Validate bug reports by creating reproducible test cases, (4) Convert issue reports into executable regression tests. Takes a repository and issue report as input and produces test code that reliably triggers the reported bug.

90

1.41x
Quality

87%

Does it follow best practices?

Impact

92%

1.41x

Average score across 3 eval scenarios

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

82%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An actionable, well-structured skill body with executable multi-language examples and a complete worked scenario. The main headroom is in trimming redundant scaffolding and promoting the run-and-verify step into an explicit workflow checkpoint.

Suggestions

Promote 'run the test to confirm it fails' from the Tips section into an explicit Step 5 validation checkpoint in the Workflow, ideally with a feedback loop (if it does not fail, refine the trigger conditions).

Trim the repeated Setup/Execute/Assert comment scaffolding duplicated across the Python, Java, and JavaScript templates, or collapse them into a single shared structure with language-specific assertion syntax only.

Consider moving the language-specific pattern templates into a references file (e.g. references/language-patterns.md) and keeping the body as a concise overview with one-level-deep pointers.

DimensionReasoningScore

Conciseness

Generally tight and well-organized with executable code, but the repeated Setup/Execute/Assert comment scaffolding across the three language templates and the Tips section (which partially restates the workflow) are minor over-explanation that could be trimmed, fitting the 4 rather than the lean 5 anchor.

4 / 5

Actionability

Provides fully executable, copy-paste-ready code for pytest, JUnit, and Jest plus a complete worked example with the exact run command and a post-fix assertion, covering the common cases comprehensively.

5 / 5

Workflow Clarity

The four-step sequence (Analyze, Inspect, Generate, Output) is clear, but validation ('Verify it fails', 'Plan for the fix') lives in the Tips section rather than as an explicit validate-checkpoint in the workflow, matching the 4 anchor with a minor validation gap; the destructive/batch cap does not apply since the skill only creates test files.

4 / 5

Progressive Disclosure

Well-organized into clearly signaled sections (Workflow, Example, Language-Specific Patterns, Constraints, Underspecified Issues, Tips) with no nested references and easy navigation, but at ~217 lines it exceeds the simple-skill exception and the language templates/Tips could be trimmed or split, so it sits at 4 rather than the cleanest 5.

4 / 5

Total

17

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, specific description that clearly conveys both capability and trigger conditions with concrete user-facing phrases. Minor headroom only in trigger-term synonym coverage.

DimensionReasoningScore

Specificity

Lists multiple concrete actions ('generates executable tests that reproduce reported bugs', 'Create a test that reproduces a bug', 'Generate failing tests from bug descriptions, stack traces, or error messages', 'Validate bug reports', 'Convert issue reports into executable regression tests'), matching the comprehensive-coverage anchor.

5 / 5

Completeness

Explicitly states what it does ('generates executable tests that reproduce reported bugs') and when to use it via a numbered 'Use when users need to' trigger list with concrete phrases, matching the both-what-and-when anchor.

5 / 5

Trigger Term Quality

Includes natural user phrases ('reproduce a bug', 'failing tests', 'stack traces', 'error messages', 'bug reports', 'regression tests') with good synonym coverage, but a few natural variations (e.g. 'write a test for this bug') are absent, so it sits below the exhaustive 5 anchor.

4 / 5

Distinctiveness Conflict Risk

Occupies a clear niche (reproduction tests from issue reports and repos) with distinct triggers and minimal overlap with other skills; uses third-person voice throughout with no first/second-person phrasing to penalize.

5 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
ArabelaTso/Skills-4-SE
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.