CtrlK
BlogDocsLog inGet started
Tessl Logo

skill-tester

Test Claude Code skills in real-world scenarios to validate functionality, usability, and effectiveness. Task-based testing operations for scenario testing, example validation, integration testing, and usability assessment. Use when testing skill functionality, validating examples work correctly, ensuring real-world effectiveness, or conducting scenario-based quality assurance.

63

Quality

75%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./.claude/skills/skill-tester/SKILL.md

The canonical home for this skill is skill-tester in fernandezbaptiste/Skrillz

SKILL.md
Quality
Evals
Security

Quality

Content

67%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is well-organized and actionable with clear validation criteria, but it restates the same four operations multiple times, adding avoidable redundancy. Strengthening error-recovery loops and de-duplicating the operation summaries would raise quality.

Suggestions

Remove the redundant restatement of the four operations: keep them in either the Overview or the Quick Reference table, not both, to tighten conciseness.

Add an explicit fix-and-retry feedback loop for failed validations (e.g., on example execution failure: diagnose, correct, re-run) to strengthen workflow_clarity.

Tighten each operation's Process steps with more specific, measurable actions (e.g., how to extract examples, what 'matches expectations' means) to lift actionability from good to fully executable.

DimensionReasoningScore

Conciseness

The four testing operations are presented three times (Overview list, per-operation detail, and the Quick Reference table), creating noticeable redundancy that could be tightened despite otherwise efficient prose.

3 / 5

Actionability

Each operation has concrete numbered steps and an explicit pass criterion ("Validation: PASS if scenario completes successfully"); as an instruction-only skill the absence of code is acceptable, with only minor gaps in specificity.

4 / 5

Workflow Clarity

Multi-step processes are clearly sequenced with PASS/FAIL validation checkpoints, though error-recovery feedback loops are weak; the operations are observational rather than destructive so the missing retry loop is a minor gap.

4 / 5

Progressive Disclosure

A single well-organized file with clear sections (Overview, When to Use, Operations, Quick Reference) and no need for external references; slightly above the 50-line simple-skill threshold but still cleanly structured.

4 / 5

Total

15

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, well-structured description that clearly states purpose and trigger conditions with natural keyword coverage. Its main weakness is that the listed actions read as testing-category labels rather than maximally concrete operations.

DimensionReasoningScore

Specificity

Lists several concrete testing operations ("scenario testing, example validation, integration testing, and usability assessment") but the actions are category labels rather than ultra-concrete verbs, leaving minor gaps in coverage.

4 / 5

Completeness

Explicitly answers both what ("Test Claude Code skills... task-based testing operations for...") and when ("Use when testing skill functionality, validating examples work correctly...") with concrete trigger phrases.

5 / 5

Trigger Term Quality

Includes natural phrases users would say ("testing skill functionality", "validating examples work correctly", "ensuring real-world effectiveness") with good coverage, though a few common synonyms are missing.

4 / 5

Distinctiveness Conflict Risk

Occupies a clear niche (testing Claude Code skills) with distinct triggers, with only minor overlap risk against closely related review/QA skills.

4 / 5

Total

17

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

Total

15

/

16

Passed

Repository
fernandezbaptiste/Skrillz
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.