Plans, generates, runs, and heals end-to-end tests using Playwright Test Agents (Planner, Generator, Healer) and the official `@playwright/mcp` server. Drives a spec-first feature-flow loop, proposes `data-testid` source diffs only when accessibility-tree locators fail, and stays token-aware via snapshot mode and `--last-failed` reruns. Use when adding E2E coverage, verifying a user journey, hardening a flaky flow, or wiring Playwright MCP into a repo. Triggers on "test this flow", "add e2e", "verify the user journey", "write e2e test", "feature test", "playwright agents", "/e2e-testing".
72
91%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
Low
Low-risk findings worth noting
Drive end-to-end tests through Playwright's MCP-backed Test Agents — Planner, Generator, Healer — released in Playwright 1.56 (Oct 2025). The user writes (or approves) a Markdown feature spec; agents generate the test, run it against a real browser via the accessibility tree, and self-heal when locators drift.
This
SKILL.mdis a thin index. Decision rules live inrules/*.mdand load on demand. Worked references (agent reference, MCP tool catalog, pyramid math) live inreferences/*.md. Literal boilerplate the skill emits lives intemplates/*.md. Do not preload everything — load only what the current phase asks for.
Reach for this skill when any of the following is true:
@playwright/mcp wiring yet and needs Phase 0 setup.Do not reach for this skill when:
tdd and the layer rule in
rules/layer-decision.md.Before any agent loop, verify the repo is wired for Playwright Test Agents. Halt and ask the user before installing anything.
Run these checks (read-only):
# 1. Playwright + MCP server installed?
jq '.devDependencies | keys[]' package.json | grep -E '@playwright/(test|mcp)'
# 2. Test-agent artefacts present?
ls specs/ tests/seed.spec.ts playwright.config.ts 2>/dev/nullDecision table:
| State | Action |
|---|---|
| Both deps present + artefacts exist | Proceed to Phase 1. |
| Deps missing | Halt. Print install plan, ask permission before running. |
| Deps present, artefacts missing | Halt. Print npx playwright init-agents --loop=claude, ask first. |
Playwright present but version < 1.56 | Halt. Test Agents require 1.56+. Ask permission to upgrade. |
Print the exact commands; do not run them silently.
The install plan template is in templates/install-plan.md.
The agent loop is spec → generate → run → heal.
The spec is human-readable Markdown, not code.
Full rules: rules/spec-first-flow.md.
specs/<flow>.md ─┐
├─→ Generator ─→ tests/<flow>.spec.ts ─→ run ─→ pass?
│ │ no
│ ▼
└──────────────────── Healer ←────────────── failing test
│
▼
patched test or `data-testid` proposalTwo entry points:
specs/<flow>.md.specs/<flow>.md.
User reviews the Markdown plan before generation.Use the Markdown template in templates/spec.md.
The Generator and the Healer both walk the accessibility tree. Pick locators in this order — never skip a rung:
getByRole('button', { name: 'Save' }) — accessibility-tree native.getByLabel, getByPlaceholder, getByText — user-facing strings.getByTestId('save-draft') — escape hatch only.data-testid is a source change, not a test workaround.
When the Healer cannot find a stable locator at rungs 1–2, propose a source
diff that adds data-testid to the component, and offer the diff for user
approval before patching the test.
Full rules and decision criteria: rules/locator-strategy.md.
Playwright MCP defaults to snapshot mode (accessibility tree, text-only).
Do not enable --caps=vision unless an explicit pixel-level concern exists.
Full rules: rules/token-budget.md.
Defaults the skill prescribes:
npx playwright test --last-failed.storageState from tests/seed.spec.ts to skip auth on every run.After the Generator produces a test:
test-provenance-guard on
the generated file to ensure the test imports production code instead
of a private re-implementation.playwright.config.ts and confirm trace: 'on-first-retry' is set
so a future failure produces a trace bundle.If the heal loop fails to converge:
confidence(analysis) on the test failure.| Signal | Do |
|---|---|
| Bug fixable by a unit or component test | Use tdd, not this skill. |
| Multi-page user flow, auth, or real network involved | Spec-first feature flow (Phase 1). |
| Flaky existing test | Healer pass only; do not rewrite without spec context. |
| Locator unstable, no stable role / label | Propose data-testid diff (rule: locator-strategy). |
| Repo missing Playwright or MCP | Phase 0 halt + ask permission. |
| Heal loop > 3 attempts | Stop, run confidence(analysis), escalate. |
| Test passes on first run, never seen failing | Run test-provenance-guard before declaring done. |
tdd — owns the unit and component layers.
This skill defers to it for anything below E2E.test-provenance-guard — runs after
Generator output to catch tests-by-construction.confidence — gate when the heal loop fails.holistic-analysis — if a flow is failing
for reasons no test rewrite can fix, step back instead of patching.playwright-trace-analyzer —
consume the trace produced by a failed test on retry.references/playwright-agents.md —
Planner / Generator / Healer reference, inputs, outputs, invocation.references/mcp-tool-catalog.md —
the @playwright/mcp tool surface, grouped by category.references/pyramid-2026.md — testing
pyramid math in 2026, with the AI-generation caveat.templates/spec.md — feature-flow Markdown spec.templates/seed.spec.ts — auth and storage
bootstrap, produces storageState.templates/playwright.config.ts —
opinionated config: snapshot mode, traces on first retry, projects per
browser, parallel CI defaults.templates/install-plan.md — Phase 0 halt
message with the exact commands to install Playwright + MCP.rules/anti-patterns.md)data-testid diff.--caps=vision without a pixel-level requirement..skip() in CI.specs/<flow>.md exists and the user reviewed it.tests/<flow>.spec.ts passes against the live app.test-provenance-guard reports no violations on the new test.playwright.config.ts has trace: 'on-first-retry'.data-testid was added, it is in the source diff and committed.39b3f44
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.