CtrlK
BlogDocsLog inGet started
Tessl Logo

skill-tester

Test Claude Code skills in real-world scenarios to validate functionality, usability, and effectiveness. Task-based testing operations for scenario testing, example validation, integration testing, and usability assessment. Use when testing skill functionality, validating examples work correctly, ensuring real-world effectiveness, or conducting scenario-based quality assurance.

60

Quality

76%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./.claude/skills/skill-tester/SKILL.md

The canonical home for this skill is skill-tester in fernandezbaptiste/Skrillz

SKILL.md
Quality
Evals
Security

Quality

Content

60%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body presents a clear four-operation structure with per-operation pass criteria, and avoids explaining concepts Claude already knows. It is weakened by redundant repetition of the same summary information in four places, generic process steps lacking concrete artifacts or measurable criteria, and no guidance on what to do when a test fails.

Suggestions

Consolidate the repeated operation summaries: the operation names and pass criteria currently appear in the Overview, 'When to Use', the Operations sections, the Quick Reference table, and the closing line — keep the Quick Reference table and trim the others.

Add concrete artifacts for the vague steps: a short PASS/FAIL test-report template for 'Document success/failure and issues', and a measurable definition of 'reasonable time' for the efficiency test.

Add failure-handling guidance: what to do when an example fails or an integration breaks (e.g., diagnose, fix, re-run, or report the failure with the skill's issue), turning 'Document any failures' into a feedback loop.

DimensionReasoningScore

Conciseness

The four operation names and pass criteria are repeated across the Overview, "When to Use", the Operations sections, the Quick Reference table, and the closing bold line ("skill-tester ensures skills work correctly through hands-on testing in real scenarios"). Mostly efficient, but the repeated summary content could be consolidated into one location.

3 / 5

Actionability

Each operation has a concrete numbered sequence ("Extract all examples from SKILL.md / Execute each example / Verify output matches expectations / Document any failures"), but key details are missing: no report template or output format for documentation, no measurable criteria for "reasonable time" in the efficiency test, and no guidance on how to select or scope scenarios.

3 / 5

Workflow Clarity

Every operation pairs a clear numbered process with an explicit validation checkpoint ("Validation: PASS if all examples work"). Scored 4 rather than 5 because error-recovery loops are absent — "Document any failures" is not a fix-retry feedback loop.

4 / 5

Progressive Disclosure

No bundle files exist, and the single SKILL.md is well-organized with clear sections (Overview, When to Use, Operations, Quick Reference) appropriate for this skill's size. Scored 4 rather than 5 because the cross-section redundancy slightly muddies navigation.

4 / 5

Total

14

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description that clearly states what the skill does, enumerates four concrete testing operations, and provides an explicit 'Use when' clause with natural trigger phrases. Its main weaknesses are mild redundancy, a few abstract buzzwords, and missing common synonyms (QA, smoke test) that would sharpen triggering.

DimensionReasoningScore

Specificity

Lists several specific actions ("scenario testing, example validation, integration testing, and usability assessment"), but the abstract claim "validate functionality, usability, and effectiveness" and partial restatement in the second sentence leave minor gaps versus fully comprehensive coverage.

4 / 5

Completeness

Explicitly answers both what ("Test Claude Code skills in real-world scenarios… Task-based testing operations for scenario testing, example validation, integration testing, and usability assessment") and when, via a "Use when…" clause with four concrete trigger phrases. Anchor 4's caveat about the 'when' being insufficiently explicit does not apply.

5 / 5

Trigger Term Quality

Natural trigger phrases like "testing skill functionality", "validating examples work correctly", and "conducting scenario-based quality assurance" are present, but common synonyms such as "QA", "smoke test", or "test report" are missing, so coverage is good rather than comprehensive.

4 / 5

Distinctiveness Conflict Risk

"Test Claude Code skills" names a clear niche with distinct triggers, but generic QA terms ("usability assessment", "scenario-based quality assurance") create minor overlap risk with broader review/QA skills.

4 / 5

Total

17

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

Total

15

/

16

Passed

Repository
fernandezbaptiste/Skrillz
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.