CtrlK
BlogDocsLog inGet started
Tessl Logo

testing-workflow

Comprehensive testing workflow orchestrating functional testing, example validation, integration testing, and usability assessment. Sequential workflow for complete skill testing from examples through scenarios to integration validation. Use when conducting thorough testing, pre-deployment validation, ensuring skill functionality, or comprehensive quality checks.

56

Quality

71%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.claude/skills/testing-workflow/SKILL.md

The canonical home for this skill is testing-workflow in fernandezbaptiste/Skrillz

SKILL.md
Quality
Evals
Security

Quality

Content

67%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a well-structured, concise orchestrator: a clear 4-step sequence with times, pass criteria, and a final validation step. Its main weakness is actionability — every step is a one-line directive with no operational detail on how to extract and run examples, define a scenario, or invoke skill-validator — and there is no failure-recovery loop despite the workflow's whole purpose being to surface failures.

Suggestions

Add concrete operational detail per step: how to extract examples from a SKILL.md, the command/form for invoking skill-validator, and what a per-step failure report should contain.

Add an explicit error-recovery loop, e.g. "If an example fails: record the failure, fix the cause, re-run all examples before proceeding to Step 2."

Trim redundancy: the Quick Reference table repeats the per-step focus/time, and the Result line appears twice — merge or cut one.

DimensionReasoningScore

Conciseness

The body is tight: one-line steps with time estimates ("**Time**: 15-30 minutes") and no explanations of concepts Claude already knows. Not 5 because the Quick Reference table repeats the step names, focuses, and times already given per-step, "Result" is stated twice, and the closing line "testing-workflow ensures skills function correctly through comprehensive hands-on testing" is filler.

4 / 5

Actionability

Some concrete guidance exists ("Test skill in 2-3 realistic scenarios", pass criteria like "All examples execute correctly"), but no executable commands or specifics on how to extract examples, execute them, or invoke skill-validator. Not 2 because step structure, counts, and pass criteria are concrete; not 4 because there is no specific, executable how-to detail for any step.

3 / 5

Workflow Clarity

A clear 4-step numbered sequence with per-step pass criteria in the Quick Reference table and a dedicated final validation step ("Run skill-validator to ensure tests didn't reveal structure issues"). Not 5 because there is no feedback loop for error recovery — failure handling ends at "❌ TESTS FAIL (with issues to fix)" with no fix-and-retry guidance; not 3 because checkpoints are explicit in the table rather than implicit.

4 / 5

Progressive Disclosure

Well-organized sections (Overview, When to Use, per-step headings, Quick Reference) and fully self-contained with no bundle files (references/scripts/assets absent), no nested or buried references. Not 5 because there are no pointers to deeper material (e.g., the component skills' operation details), and the ~72-line body exceeds the under-50-line simple-skill exception that would otherwise warrant a 5.

4 / 5

Total

15

/

20

Passed

Description

63%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description has a solid structure: an explicit what-statement, several named testing operations, and a genuine "Use when..." clause, so completeness and specificity land above the midpoint. Its weaknesses are generic, formal trigger terms that users would rarely say verbatim, and overlap risk with general testing/QA skills; it also claims "usability assessment" which the body never operationalizes.

Suggestions

Replace generic trigger phrases with natural user language, e.g. "Use when the user asks to test a skill, run its examples, QA it before deploying, or check it works after changes."

Drop "usability assessment" from the description or add it as a workflow step in the body, since the body's 4 steps never include it (an over-claim).

Lead with the distinct niche ("skill testing") and repeat it in the triggers to reduce conflict with general code-testing skills.

DimensionReasoningScore

Specificity

"orchestrating functional testing, example validation, integration testing, and usability assessment" lists several concrete, named actions in the testing domain. Not 5 because these are process labels without concrete outputs, and "usability assessment" is claimed but never defined; not 3 because more than 1-2 specific actions are listed.

4 / 5

Completeness

Explicitly answers "what" ("orchestrating functional testing, example validation, integration testing, and usability assessment") and "when" (explicit "Use when..." clause). Not 5 because the when-triggers are generic phrases rather than concrete trigger terms a user would actually say; not 3 because the when-guidance is explicitly present, not merely implied.

4 / 5

Trigger Term Quality

The trigger clause "Use when conducting thorough testing, pre-deployment validation, ensuring skill functionality, or comprehensive quality checks" contains relevant keywords but they are formal and generic, missing natural user phrasings like "test this skill", "run tests", or "QA the skill". Not 4 because common natural variations are absent; not 2 because multiple relevant trigger phrases are present.

3 / 5

Distinctiveness Conflict Risk

"complete skill testing" and "ensuring skill functionality" anchor it to the skill-testing niche, but "conducting thorough testing" and "comprehensive quality checks" would fire for virtually any testing or QA request, creating real overlap risk with general testing skills. Not 4 because the generic quality-check triggers carry more than minor overlap risk.

3 / 5

Total

14

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

Total

15

/

16

Passed

Repository
fernandezbaptiste/Skrillz
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.