CtrlK
BlogDocsLog inGet started
Tessl Logo

e2e-testing

Plans, generates, runs, and heals end-to-end tests using Playwright Test Agents (Planner, Generator, Healer) and the official `@playwright/mcp` server. Drives a spec-first feature-flow loop, proposes `data-testid` source diffs only when accessibility-tree locators fail, and stays token-aware via snapshot mode and `--last-failed` reruns. Use when adding E2E coverage, verifying a user journey, hardening a flaky flow, or wiring Playwright MCP into a repo. Triggers on "test this flow", "add e2e", "verify the user journey", "write e2e test", "feature test", "playwright agents", "/e2e-testing".

72

Quality

91%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

E2E Testing

Drive end-to-end tests through Playwright's MCP-backed Test Agents — Planner, Generator, Healer — released in Playwright 1.56 (Oct 2025). The user writes (or approves) a Markdown feature spec; agents generate the test, run it against a real browser via the accessibility tree, and self-heal when locators drift.

This SKILL.md is a thin index. Decision rules live in rules/*.md and load on demand. Worked references (agent reference, MCP tool catalog, pyramid math) live in references/*.md. Literal boilerplate the skill emits lives in templates/*.md. Do not preload everything — load only what the current phase asks for.


When to use

Reach for this skill when any of the following is true:

  • A feature has user-facing flow that integration tests cannot fully cover.
  • A bug repros only through real navigation (multi-page, auth, real network).
  • A flake needs a Healer pass instead of a manual locator hunt.
  • The repo has no @playwright/mcp wiring yet and needs Phase 0 setup.

Do not reach for this skill when:

  • A unit or component test would catch the same bug — defer to tdd and the layer rule in rules/layer-decision.md.
  • The change is a pure refactor with no behavioural surface.
  • You are adding test infrastructure unrelated to a real flow.

Phase 0 — Preflight (mandatory gate)

Before any agent loop, verify the repo is wired for Playwright Test Agents. Halt and ask the user before installing anything.

Run these checks (read-only):

# 1. Playwright + MCP server installed?
jq '.devDependencies | keys[]' package.json | grep -E '@playwright/(test|mcp)'

# 2. Test-agent artefacts present?
ls specs/ tests/seed.spec.ts playwright.config.ts 2>/dev/null

Decision table:

StateAction
Both deps present + artefacts existProceed to Phase 1.
Deps missingHalt. Print install plan, ask permission before running.
Deps present, artefacts missingHalt. Print npx playwright init-agents --loop=claude, ask first.
Playwright present but version < 1.56Halt. Test Agents require 1.56+. Ask permission to upgrade.

Print the exact commands; do not run them silently. The install plan template is in templates/install-plan.md.


Phase 1 — Spec-first feature flow

The agent loop is spec → generate → run → heal. The spec is human-readable Markdown, not code. Full rules: rules/spec-first-flow.md.

specs/<flow>.md   ─┐
                   ├─→  Generator  ─→  tests/<flow>.spec.ts  ─→  run  ─→  pass?
                   │                                                       │ no
                   │                                                       ▼
                   └────────────────────  Healer  ←──────────────  failing test
                                              │
                                              ▼
                                  patched test or `data-testid` proposal

Two entry points:

  1. Spec already drafted by the user. Skip the Planner. Run the Generator on specs/<flow>.md.
  2. App exists, no spec yet. Run the Planner against the live app to draft specs/<flow>.md. User reviews the Markdown plan before generation.

Use the Markdown template in templates/spec.md.

Locator ladder (when generating or healing)

The Generator and the Healer both walk the accessibility tree. Pick locators in this order — never skip a rung:

  1. getByRole('button', { name: 'Save' }) — accessibility-tree native.
  2. getByLabel, getByPlaceholder, getByText — user-facing strings.
  3. getByTestId('save-draft') — escape hatch only.

data-testid is a source change, not a test workaround. When the Healer cannot find a stable locator at rungs 1–2, propose a source diff that adds data-testid to the component, and offer the diff for user approval before patching the test. Full rules and decision criteria: rules/locator-strategy.md.


Phase 2 — Token-aware execution

Playwright MCP defaults to snapshot mode (accessibility tree, text-only). Do not enable --caps=vision unless an explicit pixel-level concern exists. Full rules: rules/token-budget.md.

Defaults the skill prescribes:

  • Snapshot mode (no vision) for all agent calls.
  • Run only the changed spec on iteration: npx playwright test --last-failed.
  • Run the Healer only on failure, not on every save.
  • Reuse storageState from tests/seed.spec.ts to skip auth on every run.
  • Cap the heal loop at three attempts per failing test before escalating.

Phase 3 — Verification

After the Generator produces a test:

  1. Run the test once against the live app. It must pass on first run, or the Healer must converge in ≤ 3 attempts.
  2. Invoke test-provenance-guard on the generated file to ensure the test imports production code instead of a private re-implementation.
  3. Open playwright.config.ts and confirm trace: 'on-first-retry' is set so a future failure produces a trace bundle.

If the heal loop fails to converge:

  • Invoke confidence(analysis) on the test failure.
  • If confidence is below 90%, escalate to the user with the trace, the spec, and the proposed locator changes — do not keep healing blindly.

Decision flow at a glance

SignalDo
Bug fixable by a unit or component testUse tdd, not this skill.
Multi-page user flow, auth, or real network involvedSpec-first feature flow (Phase 1).
Flaky existing testHealer pass only; do not rewrite without spec context.
Locator unstable, no stable role / labelPropose data-testid diff (rule: locator-strategy).
Repo missing Playwright or MCPPhase 0 halt + ask permission.
Heal loop > 3 attemptsStop, run confidence(analysis), escalate.
Test passes on first run, never seen failingRun test-provenance-guard before declaring done.

Composes with

  • tdd — owns the unit and component layers. This skill defers to it for anything below E2E.
  • test-provenance-guard — runs after Generator output to catch tests-by-construction.
  • confidence — gate when the heal loop fails.
  • holistic-analysis — if a flow is failing for reasons no test rewrite can fix, step back instead of patching.
  • playwright-trace-analyzer — consume the trace produced by a failed test on retry.

References

Templates


Anti-patterns (one-liner — full list in rules/anti-patterns.md)

  • Writing E2E for logic a unit test catches.
  • Running the Healer on every save.
  • Patching the test with brittle CSS selectors instead of proposing a data-testid diff.
  • Enabling --caps=vision without a pixel-level requirement.
  • Ignoring Healer suggestions and keeping a .skip() in CI.
  • Generating tests against a stub server, not the real app.

Definition of done

  • Phase 0 preflight passed or installs were user-approved.
  • specs/<flow>.md exists and the user reviewed it.
  • tests/<flow>.spec.ts passes against the live app.
  • test-provenance-guard reports no violations on the new test.
  • playwright.config.ts has trace: 'on-first-retry'.
  • If a data-testid was added, it is in the source diff and committed.
Repository
mthines/agent-skills
Last updated
First committed

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.