CtrlK
BlogDocsLog inGet started
Tessl Logo

qa

Autonomous multi-app QA sweep that drives template apps with Playwright MCP. Use only when the user explicitly runs /qa or asks for an end-to-end QA sweep. Do not load it for ordinary feature work, single-page checks, or one-off bug fixes.

65

Quality

82%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

77%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is an unusually actionable, well-sequenced runbook: every step has executable commands, validation checkpoints, failure fallbacks, and bounded loops. Its weaknesses are token efficiency (over-explained Playwright basics, an inline 115-line prompt template) and the total absence of progressive disclosure — everything is monolithically inlined with no reference files.

Suggestions

Move the full tester prompt template into a bundle file (e.g. references/tester-prompt.md) and keep only the placeholder list and a one-line pointer in SKILL.md, cutting ~100 lines from the always-loaded context.

Replace the nine-item Playwright tool walkthrough with a compact table or one sentence ('use browser_navigate / browser_snapshot / browser_click / browser_type / browser_fill_form; check browser_console_messages and browser_network_requests after each interaction') since Claude already knows these tools.

Merge overlapping guidance to tighten conciseness: the credential-handling rules in Step 2 duplicate instructions the tester prompt already carries, and the 'What Counts as a Bug' list could be folded into the report-format section of the offloaded tester prompt.

DimensionReasoningScore

Conciseness

Most of the body is operational detail Claude could not infer (ports, per-app credential table, HMR timings, troubleshooting causes), but it is padded by over-explanation of things Claude already knows — the Playwright MCP tool list ('See the page: browser_snapshot', 'Click: browser_click', 'Type: browser_type') and a ~115-line inline tester prompt template. This matches 'mostly efficient but includes some unnecessary explanation or could be tightened' rather than 4, whose trims would be minor.

3 / 5

Actionability

Guidance is fully executable: exact dev-server commands with ports ('cd templates/mail && PORT=9201 pnpm dev'), a copy-paste readiness poll, a port-kill command ('lsof -ti :9201 | xargs kill -9'), a complete fill-in-the-blanks tester prompt, and a fixed report format. This matches the top anchor of copy-paste ready commands covering the common cases.

5 / 5

Workflow Clarity

The orchestrator steps are numbered 1-7 with explicit validation checkpoints (curl readiness polling with a 30s timeout and skip-and-report fallback), and the tester loop includes feedback loops: fix, wait for HMR, retest, escalate after 3 attempts, 'Maximum 2 full passes. Don't loop forever', plus a final regression pass. Not a 4 because no validation step is missing or implicit.

5 / 5

Progressive Disclosure

No bundle files exist and all content lives in one ~320-line SKILL.md, including a ~115-line tester prompt template that clearly belongs in a separate reference file. Section headers give it real structure (better than the headerless 300-line anchor of a 2), but there are no references at all and inlineable-offloadable content remains, matching 'some structure but could be better organized'.

3 / 5

Total

16

/

20

Passed

Description

82%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is strong on completeness and distinctiveness, with an explicit 'use when' trigger, concrete /qa invocation terms, and negative guidance that prevents conflicts with routine verification work. The only soft spot is specificity: the description omits the fix-retest-report cycle that the body delivers.

DimensionReasoningScore

Specificity

The description names the domain ('Autonomous multi-app QA sweep') and one concrete action ('drives template apps with Playwright MCP'), but the fix/retest/report cycle described in the body is absent, so action coverage is not comprehensive. It sits at the 'domain plus 1-2 concrete actions' anchor rather than 4, which requires several specific actions.

3 / 5

Completeness

Both halves are explicit: the what ('Autonomous multi-app QA sweep that drives template apps with Playwright MCP') and the when with concrete trigger phrases ('Use only when the user explicitly runs /qa or asks for an end-to-end QA sweep'). This is a clean match for the top anchor; a 4 would require the 'when' to be less explicit.

5 / 5

Trigger Term Quality

Natural terms users would say are present: 'QA sweep', 'end-to-end QA sweep', '/qa', 'Playwright', 'bug fixes'. A few common variations are missing (e.g., 'e2e', 'regression test', 'test the apps'), which keeps it below the comprehensive synonym coverage of a 5.

4 / 5

Distinctiveness Conflict Risk

The trigger is tightly scoped ('Use only when the user explicitly runs /qa') and reinforced with negative triggers ('Do not load it for ordinary feature work, single-page checks, or one-off bug fixes'), which cleanly separates it from general verification or code-review skills. Minimal conflict risk.

5 / 5

Total

17

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

13

/

16

Passed

Repository
BuilderIO/agent-native
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.