CtrlK
BlogDocsLog inGet started
Tessl Logo

e2e-pr-stabilizer

Stabilizes or optimizes Playwright E2E tests on one PR via a local-first loop, then ratifies with a single CI run. Pulls Dash0 spans (`git.pull_request_link`) as the historical baseline, then captures every iteration's evidence locally with `--trace=on` (same OTel exporter, same trace schema). Validation is empirical, not predictive: before commit, every new locator must resolve against source (static grep) or the live app (`locator.count()`); after commit, the fixed test must pass three consecutive local runs before the single push. Modes: `stabilize` (default) heals flaky / failing tests; `optimize` is report-only and ranks slow-action wins by measured ms saved. Refuses `.skip`, `.fixme`, `waitForTimeout`, or any check-weakening edit. Use when a PR has flaky or failing E2E tests or when you want to find slow tests worth tightening. Triggers on "stabilize this PR", "fix flaky e2e", "heal playwright on PR", "ui-e2e is failing", "self-heal e2e", "optimize e2e", "/e2e-pr-stabilizer".

64

Quality

81%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

58%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is an unusually well-sequenced orchestrator document — explicit phases, gates, thresholds, feedback loops, and per-mode checklists — with workflow clarity at the top anchor. Its two real weaknesses are that the load-on-demand architecture points at bundle files that are not actually shipped, making the detailed procedures unreachable, and that core rules are repeated many times over, inflating token cost without adding information.

Suggestions

Ship the referenced rules/*.md and templates/*.md files in the bundle (or inline the essential per-phase procedures into SKILL.md) — 9 of 10 referenced paths currently do not exist, so the 'load on demand' design dead-ends at every phase.

State each invariant once: the three-consecutive-passes rule and the guard-rail ban list each appear five to seven times across the intro, Modes table, Workflow table, Core Principles, Quickstart, Definition of Done, and Anti-patterns; consolidate them into the Workflow/gate column and rules/guard-rails.md.

Repair broken links inside the bundle: references/dash0-mcp-filters.md points to ../rules/telemetry-driven-analysis.md which is missing, and the ../../ cross-skill and ../../../agents/shared paths are unverifiable from this bundle.

DimensionReasoningScore

Conciseness

The body is dense and skill-specific with no generic-concept padding, but key facts are heavily repeated — the three-consecutive-local-passes rule is restated roughly seven times ("never commits a fix until three consecutive local runs prove it works", "passes 3 times in a row", "Three consecutive local passes or no commit", Quickstart step 6, the Definition of Done, and two anti-patterns) and the `.skip`/`.fixme`/`waitForTimeout` ban appears five times; the Modes and Core Principles tables also restate Workflow-table content. Anchor 3 ("mostly efficient but could be tightened") fits: more than minor trimming is needed, but nothing explains concepts Claude already knows, so it is above 2.

3 / 5

Actionability

There is genuinely concrete guidance inline ("failure_rate ≥ 0.10 over ≥ 5 attempts, or flake_count ≥ 2", "dur > 5×median", "--trace=on", "--repeat-each=3", "locator.count() ≥ 1", "npx playwright init-agents --loop=claude", and copy-ready slash invocations), but the per-phase executable procedures are delegated to rules/*.md files that do not exist in the bundle — only references/dash0-mcp-filters.md is present — so an agent cannot actually execute Phases 0–8 from what ships. That is "some concrete guidance but incomplete; missing key details" (3), not 4, because the missing delegation targets break executability rather than leaving minor gaps.

3 / 5

Workflow Clarity

Eight phases are explicitly sequenced in a table with a named rule file and an explicit gate per phase, validation is deterministic and front-loaded (Phase 5 locator-existence check, Phase 6's "passes 3 times in a row" streak with "a single failure or flake within the streak resets the counter" and "maximum 10 attempts per test before escalating"), and error-recovery feedback loops are explicit ("discard the diff and re-enter Phase 4 with that evidence"; "If CI disagrees with the local result, that is a signal to escalate, not to re-enter the loop blindly"). Definition-of-Done checklists per mode complete the match for anchor 5.

5 / 5

Progressive Disclosure

The thin-index design is well conceived — per-phase "Load on demand. Do not preload" tables, clearly signaled one-level-deep references — but scored against the actual bundle, 9 of the 10 referenced files (all eight rules/*.md plus templates/stabilization-report.md) are missing; only references/dash0-mcp-filters.md exists, and it itself links to the missing ../rules/telemetry-driven-analysis.md, while cross-skill links (../../analysis/playwright-trace-analyzer, ../../delivery/ci-auto-fix, ../../../agents/shared/...) are also unverifiable. Navigation dead-ends at most references, which lands below the midpoint (2) rather than at 3, where the issue would merely be unclear signaling or inline bloat.

2 / 5

Total

13

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is exemplary: third-person, concrete about both modes' capabilities, explicit about what it refuses to do, and equipped with an explicit 'Use when' clause plus seven literal trigger phrases including the slash command. It clearly fits the top anchor on all four dimensions with no vagueness or over-claiming.

DimensionReasoningScore

Specificity

Multiple concrete actions with comprehensive coverage across both modes: "Stabilizes or optimizes Playwright E2E tests on one PR", "Pulls Dash0 spans (git.pull_request_link) as the historical baseline", "every new locator must resolve against source (static grep) or the live app (locator.count())", "must pass three consecutive local runs before the single push", "ranks slow-action wins by measured ms saved", "Refuses .skip, .fixme, waitForTimeout". Consistent third person ("Stabilizes", "Pulls", "Refuses"). Not 4: there are no gaps in action coverage — telemetry, local iteration, validation, modes, and refusal policy are all named concretely.

5 / 5

Completeness

Explicitly answers both questions: the "what" is concrete (stabilize/optimize Playwright E2E tests on one PR via a local-first loop, telemetry baseline, locator validation, 3-pass gate, single CI ratification) and the "when" is an explicit clause with concrete trigger phrases ("Use when a PR has flaky or failing E2E tests or when you want to find slow tests worth tightening"). This matches the anchor-5 example pattern exactly; a missing or implied 'Use when' would cap it at 3, but it is present and specific.

5 / 5

Trigger Term Quality

Comprehensive natural-term coverage including synonyms and the slash command: "stabilize this PR", "fix flaky e2e", "heal playwright on PR", "ui-e2e is failing", "self-heal e2e", "optimize e2e", "/e2e-pr-stabilizer", plus "Use when a PR has flaky or failing E2E tests or when you want to find slow tests worth tightening". Phrases like "fix flaky e2e" and "ui-e2e is failing" are exactly what a user would say; not 4 because the synonym set (flaky / failing / heal / self-heal / optimize / stabilize) and the literal command invocation leave few natural variants out.

5 / 5

Distinctiveness Conflict Risk

A clear niche with distinct triggers: Playwright E2E flake/slow-test stabilization scoped to a single PR, anchored on Dash0 telemetry, trace evidence, and a 3-consecutive-local-pass gate — vocabulary no generic CI-fix or test-writing skill shares. Not 4: even against a closely related skill like a generic ci-auto-fix, the Playwright/E2E/flake framing and named trigger phrases keep conflict risk minimal.

5 / 5

Total

20

/

20

Passed

Validation

75%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 12 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

relative_links

Relative link issues: 31 missing, 5 suspicious

Warning

Total

12

/

16

Passed

Repository
mthines/agent-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.