CtrlK
BlogDocsLog inGet started
Tessl Logo

testing-validator

Comprehensive testing validation for Claude Code skills through functional testing, example validation, integration testing, regression testing, and edge case testing. Task-based testing operations with automated example execution, manual scenario testing, and test reporting. Use when testing skill functionality, validating examples execute correctly, ensuring integration works, preventing regressions, or conducting comprehensive functional quality assurance.

58

Quality

73%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/testing-validator/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

48%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is well-structured with clear operations, checklists, and pass criteria, but it is severely padded with fabricated example outputs, duplicated sections, and generic testing explanations. The automation commands reference scripts and guide files that are absent from the bundle, undermining both actionability and progressive disclosure.

Suggestions

Cut the fabricated sample-output example blocks (Operations 1-5) and move condensed versions into the referenced guides, keeping only the process, checklist, and pass criteria inline.

Actually provide the referenced 'scripts/validate-examples.py', 'test-runner.py', 'generate-test-report.py', and the six references/*.md files — or remove those references — since the skill's automation story depends entirely on files that don't exist.

Merge the overlapping 'Best Practices' and 'Common Mistakes' sections and drop the duplicated review-multi comparison and operations table from the Quick Reference to roughly halve the token footprint.

DimensionReasoningScore

Conciseness

The ~830-line body is noticeably verbose: roughly 250 lines are fabricated sample-output blocks (e.g., the 35-line 'Functional Testing: skill-researcher' example), the Quick Reference section duplicates the operations table, Best Practices and Common Mistakes restate each other, and generic QA concepts (edge cases like empty or invalid inputs) are explained at length despite being knowledge Claude already has.

2 / 5

Actionability

There are concrete commands ('python3 scripts/validate-examples.py /path/to/skill', 'test-runner.py --mode comprehensive') and quantified pass criteria (≥90% success rate), but the referenced scripts/ and references/ files do not exist in the bundle, and most process steps are abstract direction like 'Actually follow skill instructions' rather than executable guidance.

3 / 5

Workflow Clarity

Each operation has numbered steps, a validation checklist, explicit PASS/PARTIAL/FAIL criteria with thresholds, and the regression operation defines a baseline → change → re-run → compare feedback loop. It falls short of 5 because the Comprehensive mode's 'Aggregate results / Make deployment decision' steps are undefined and the aggregation logic is left implicit.

4 / 5

Progressive Disclosure

References are clearly signaled in a 'For More Information' section, but none of the six referenced files or three scripts actually exist in the bundle, and the in-depth operation detail and example blocks that belong in those files are inlined, making the ~830-line body effectively monolithic.

3 / 5

Total

12

/

20

Passed

Description

88%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description that clearly and explicitly states both what the skill does and when to use it, enumerating all five testing operations as concrete actions. Its only weaknesses are repetitive 'testing' vocabulary that omits common synonyms and slight overlap risk with general testing requests.

DimensionReasoningScore

Specificity

The description lists multiple concrete actions covering all five operations ('functional testing, example validation, integration testing, regression testing, and edge case testing') plus 'automated example execution, manual scenario testing, and test reporting', giving comprehensive coverage with no gaps.

5 / 5

Completeness

It explicitly answers both what ('Comprehensive testing validation for Claude Code skills through... Task-based testing operations with automated example execution...') and when ('Use when testing skill functionality, validating examples execute correctly, ensuring integration works, preventing regressions...') with concrete trigger phrases.

5 / 5

Trigger Term Quality

'Use when testing skill functionality, validating examples execute correctly, ensuring integration works, preventing regressions' are natural user phrases, but the vocabulary leans almost entirely on 'testing' and misses common synonyms such as 'does this skill work', 'check the skill', or 'run the tests'.

4 / 5

Distinctiveness Conflict Risk

The niche is distinct ('for Claude Code skills'), but broadly common terms like 'testing', 'regressions', and 'quality assurance' create minor overlap risk with ordinary code-testing requests.

4 / 5

Total

18

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (832 lines); consider splitting into references/ and linking

Warning

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

referenced_paths_exist

Referenced path issues: 14 missing

Warning

Total

13

/

16

Passed

Repository
fernandezbaptiste/Skrillz
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.