CtrlK
BlogDocsLog inGet started
Tessl Logo

testing-validator

Comprehensive testing validation for Claude Code skills through functional testing, example validation, integration testing, regression testing, and edge case testing. Task-based testing operations with automated example execution, manual scenario testing, and test reporting. Use when testing skill functionality, validating examples execute correctly, ensuring integration works, preventing regressions, or conducting comprehensive functional quality assurance.

55

Quality

69%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.claude/skills/testing-validator/SKILL.md

The canonical home for this skill is testing-validator in fernandezbaptiste/Skrillz

SKILL.md
Quality
Evals
Security

Quality

Content

45%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is well-structured for workflow clarity — each operation has sequenced steps, checklists, and explicit pass criteria — but it is heavily padded with ~300 lines of illustrative test reports and its entire automation and reference layer is broken: every referenced script and reference file is missing from the bundle. The skill reads as a process framework rather than an immediately executable guide.

Suggestions

Create the referenced bundle files (scripts/validate-examples.py, scripts/test-runner.py, scripts/generate-test-report.py, and the six references/*-guide.md files) or remove all references to them — every referenced path currently resolves to a nonexistent file.

Replace the five ~50-line illustrative "Example" report blocks with short 5-10 line templates (or move one full sample to references/test-report-template.md), cutting roughly 300 lines of padding.

De-duplicate the triple restatement of the five operations: keep the Quick Reference table and per-operation detail, and trim the Overview bullet list and closing summary line that repeat the same information.

DimensionReasoningScore

Conciseness

The ~830-line body is noticeably verbose: five full-page illustrative "Example" blocks (~300 lines total) are sample output reports of other skills (skill-researcher, review-multi, todo-management, development-workflow) with limited instructional value, and the 5 operations are restated three times (Overview bullets, full Operations sections, Quick Reference tables) plus a closing summary line. Not a 1 because the material is structured process guidance rather than explanations of concepts Claude already knows.

2 / 5

Actionability

Concrete commands are present ("python3 scripts/validate-examples.py /path/to/skill", "python3 scripts/test-runner.py /path/to/skill --mode comprehensive") but they are broken as written — the bundle contains no scripts/ or references/ directories, so the primary automation guidance cannot execute. The manual guidance (per-operation 5-step processes and checklists) is reasonably concrete, which keeps this above the minimal-guidance anchor at 2.

3 / 5

Workflow Clarity

Every operation has a clearly sequenced 5-step process, an explicit Validation Checklist, PASS/PARTIAL/FAIL decision criteria, and time estimates; regression testing includes an inherent re-run/compare feedback loop. It stops short of 5 because there is no explicit gate for what to do when an operation fails mid-comprehensive-run (abort, skip, or continue) and the Comprehensive Mode's aggregate/decision step is thin.

4 / 5

Progressive Disclosure

The "For More Information" section signals one-level-deep references (references/functional-testing-guide.md, example-validation-guide.md, etc.) and the Automation Scripts section points to three scripts — but none of these files exist in the bundle, so navigation leads nowhere. Meanwhile ~300 lines of sample test reports that clearly belong in a separate template file are inlined in SKILL.md. This matches the anchor of content belonging in separate files being inlined, with references effectively broken.

2 / 5

Total

11

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description that explicitly states what the skill does (five concrete testing operations) and when to use it, with natural trigger phrases and a clear skill-testing niche. Its only weaknesses are buzzword padding ("comprehensive" appears twice) and slight overlap risk with general QA and companion review skills.

DimensionReasoningScore

Specificity

The description lists several concrete actions — "functional testing, example validation, integration testing, regression testing, and edge case testing" plus "automated example execution, manual scenario testing, and test reporting" — giving broad coverage. It falls just short of the 5 anchor because the opening "Comprehensive testing validation" and closing "comprehensive functional quality assurance" are vague buzzword padding that repeats sentence one's content rather than adding capability.

4 / 5

Completeness

Both questions are explicitly answered: the "what" is the enumerated testing operations, and the "when" is an explicit clause with concrete triggers ("Use when testing skill functionality, validating examples execute correctly, ensuring integration works, preventing regressions, or conducting comprehensive functional quality assurance"). This matches the anchor for a clear and explicit what-plus-when with concrete trigger phrases.

5 / 5

Trigger Term Quality

The "Use when testing skill functionality, validating examples execute correctly, ensuring integration works, preventing regressions" clause contains multiple natural phrases a user would say. A few common variations are missing (e.g., "QA", "verify the skill works", "smoke test"), keeping it below the comprehensive-synonyms anchor at 5.

4 / 5

Distinctiveness Conflict Risk

The scope is pinned to "Claude Code skills" testing, a clear niche with distinct triggers. Minor overlap risk remains with closely related skills: general code-testing/QA skills could trigger on "testing" and "quality assurance", and the companion review-multi skill (per the body) shares the "comprehensive quality" trigger space.

4 / 5

Total

17

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (832 lines); consider splitting into references/ and linking

Warning

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

referenced_paths_exist

Referenced path issues: 14 missing

Warning

Total

13

/

16

Passed

Repository
fernandezbaptiste/Skrillz
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.