CtrlK
BlogDocsLog inGet started
Tessl Logo

skill-tester

Test Claude Code skills in real-world scenarios to validate functionality, usability, and effectiveness. Task-based testing operations for scenario testing, example validation, integration testing, and usability assessment. Use when testing skill functionality, validating examples work correctly, ensuring real-world effectiveness, or conducting scenario-based quality assurance.

60

Quality

76%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./skills/skill-tester/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

60%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is cleanly organized and each of the four testing operations has a defined sequence with pass criteria, but the same content is repeated four times and the steps stay abstract — there is no concrete guidance on how to execute or verify examples, nor any failure-handling feedback loop. As an instruction-only skill it is serviceable but would leave a reader guessing at execution details.

Suggestions

Collapse the duplication: keep the Operations sections and one summary table, and drop the repeated four-item lists in Overview, When to Use, and the closing tagline.

Make steps executable: specify how to extract and run examples from a target SKILL.md (e.g., identifying code blocks vs. prose instructions, sandboxing risky commands) and what a documented failure report must contain.

Add a feedback loop for failures: after a failed example or integration test, instruct the model to diagnose the cause, note whether it is a skill defect, and re-run to confirm the outcome.

DimensionReasoningScore

Conciseness

The same four operations are restated four times — the Overview list, When to Use, the Operations sections, and the Quick Reference table — plus a closing tagline and a "Use with" line that repeats the review-multi integration note already given in Overview. It never explains concepts Claude already knows, so it is not anchor-2 verbose, but the systematic duplication goes beyond the minor trimming of anchor 4.

3 / 5

Actionability

Each operation has a numbered process and a pass criterion, but the steps are abstract directives like "Execute each example" and "Verify output matches expectations" with no specifics on how to execute examples, what counts as a match, or what a failure report should contain. This is incomplete concrete guidance rather than the mostly-executable level of anchor 4, and well above the high-level-hints-only level of anchor 2.

3 / 5

Workflow Clarity

Every operation has a clearly sequenced numbered process capped by an explicit "Validation: PASS if..." checkpoint, which matches a clear sequence with most checkpoints present. It is not anchor 5 because there are no feedback loops — no guidance on what to do when an example or integration fails beyond "document any failures".

4 / 5

Progressive Disclosure

The body is a single well-organized overview level with clear section headers and no external file references to go stale; no bundle files (references/, scripts/, assets/) exist, and the body references none. It is not anchor 5 because the body exceeds the simple-skill line and carries duplicated content (Overview, When to Use, Quick Reference, tagline) that should be consolidated.

4 / 5

Total

14

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is strong: it names four concrete testing operations, uses third person, and pairs a clear 'what' with an explicit 'Use when' trigger clause. Its only weaknesses are mild abstraction in the opening sentence, missing a few natural trigger synonyms, and slight overlap risk with general software-testing vocabulary.

DimensionReasoningScore

Specificity

"Task-based testing operations for scenario testing, example validation, integration testing, and usability assessment" lists several concrete operations with minor gaps, matching the anchor for several specific actions. Not 5 because "validate functionality, usability, and effectiveness" remains abstract; not 3 because four distinct operations are named. Third person voice is used throughout, so no voice penalty applies.

4 / 5

Completeness

It clearly answers both what ("Test Claude Code skills in real-world scenarios... Task-based testing operations for scenario testing, example validation, integration testing, and usability assessment") and when ("Use when testing skill functionality, validating examples work correctly, ensuring real-world effectiveness, or conducting scenario-based quality assurance") with concrete trigger phrases. Matches the top anchor exactly.

5 / 5

Trigger Term Quality

The "Use when" clause contains natural user phrases like "testing skill functionality", "validating examples work correctly", and "scenario-based quality assurance", giving good keyword coverage. A few natural synonyms (e.g., smoke test, dry run, QA a skill) are missing, so it falls just below the comprehensive anchor.

4 / 5

Distinctiveness Conflict Risk

The niche — testing Claude Code skills — is clearly distinct, with triggers tied to skill validation. Minor overlap risk remains with general software-testing requests since terms like "integration testing" could fire for ordinary test-writing asks, placing it just below the clear-niche anchor.

4 / 5

Total

17

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

Total

15

/

16

Passed

Repository
fernandezbaptiste/Skrillz
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.