CtrlK
BlogDocsLog inGet started
Tessl Logo

e2e-testing

Plans, generates, runs, and heals end-to-end tests using Playwright Test Agents (Planner, Generator, Healer) and the official `@playwright/mcp` server. Drives a spec-first feature-flow loop, proposes `data-testid` source diffs only when accessibility-tree locators fail, and stays token-aware via snapshot mode and `--last-failed` reruns. Use when adding E2E coverage, verifying a user journey, hardening a flaky flow, or wiring Playwright MCP into a repo. Triggers on "test this flow", "add e2e", "verify the user journey", "write e2e test", "feature test", "playwright agents", "/e2e-testing".

72

Quality

91%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

81%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-engineered thin-index skill: phased workflow with explicit gates and escalation loops, executable commands, and disciplined token economy. The main gaps are minor — a small amount of trimmable time-sensitive detail and heavy reliance on rules/templates files that are not present in the provided bundle.

Suggestions

Inline the exact Planner/Generator/Healer invocation commands (or a single one-line example each) so the core loop is executable from SKILL.md alone.

Move the 'released in Playwright 1.56 (Oct 2025)' date detail out of the intro (the Phase 0 version gate already covers the requirement functionally), keeping time-sensitive facts in a dedicated section.

Ensure the referenced rules/*.md and templates/* files ship with the skill bundle — 9 of the 12 referenced paths are missing from the provided file listing, which breaks the thin-index navigation the skill depends on.

DimensionReasoningScore

Conciseness

The body is dominated by terse decision tables, a compact locator ladder with real code, and a one-line anti-patterns list — efficient overall. It is not 5 because of minor trimmable detail, e.g. 'released in Playwright 1.56 (Oct 2025)' is time-sensitive information outside a deprecation/old-patterns section, and the Phase 0 prose slightly restates the decision table.

4 / 5

Actionability

Most guidance is executable as written: the jq/grep preflight check, 'npx playwright init-agents --loop=claude', 'npx playwright test --last-failed', "getByRole('button', { name: 'Save' })", and the 'trace: on-first-retry' config check. It is not 5 because the actual Planner/Generator/Healer invocation commands are deferred to references/playwright-agents.md rather than being copy-paste ready in the body.

4 / 5

Workflow Clarity

The Phase 0→3 sequence is explicit with a mandatory preflight gate, decision tables for routing, a heal loop capped at three attempts with a confidence-threshold escalation feedback loop, and a Definition-of-done checklist containing verification steps (first-run pass, provenance guard, trace config). Checkpoints and error-recovery loops are all present.

5 / 5

Progressive Disclosure

The body declares itself a thin index ('Do not preload everything — load only what the current phase asks for') with one-level-deep, phase-labeled references, and the three referenced references/*.md files all exist in the bundle. It is not 5 because the 5 referenced rules/*.md and 4 templates/* files are absent from the provided bundle, so most referenced paths cannot be verified to resolve.

4 / 5

Total

17

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A model description: concrete capability list, third-person voice, explicit use-when clause, and natural trigger phrases with synonyms and the slash command. It is specific enough to avoid conflicts with adjacent testing skills while remaining concise.

DimensionReasoningScore

Specificity

The description enumerates multiple concrete actions — 'Plans, generates, runs, and heals end-to-end tests', 'proposes data-testid source diffs only when accessibility-tree locators fail', 'snapshot mode and --last-failed reruns' — giving comprehensive, non-generic coverage of the skill's capabilities. It is not 4 because there are no meaningful gaps in the action inventory.

5 / 5

Completeness

It explicitly answers both 'what' (spec-first plan/generate/run/heal loop via Planner, Generator, Healer and @playwright/mcp) and 'when' ('Use when adding E2E coverage, verifying a user journey, hardening a flaky flow, or wiring Playwright MCP into a repo') with concrete trigger phrases. A missing 'Use when' clause would cap this at 3; here the clause is explicit.

5 / 5

Trigger Term Quality

Trigger phrases are natural user utterances with synonym coverage: 'test this flow', 'add e2e', 'verify the user journey', 'write e2e test', 'feature test', 'playwright agents', '/e2e-testing', plus 'hardening a flaky flow' and 'wiring Playwright MCP'. It is not 4 because no commonly used variant of the request is obviously missing.

5 / 5

Distinctiveness Conflict Risk

The niche is clear — Playwright Test Agents with the official MCP server for E2E flows — and triggers like 'playwright agents' and '/e2e-testing' are unlikely to fire for unit-test or general-testing skills. Voice is third person ('Plans, generates…'), so no specificity penalty applies.

5 / 5

Total

20

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_field

'metadata' should map string keys to string values

Warning

relative_links

Relative link issues: 13 missing, 8 suspicious

Warning

Total

14

/

16

Passed

Repository
mthines/agent-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.