CtrlK
BlogDocsLog inGet started
Tessl Logo

debug-e2e-test

Use when investigating a specific named Positron e2e (Playwright) test that is failing or flaking -- in CI ("why is <test> failing on main", "this test is flaky in CI", "is this a test bug or a product bug") or on the engineer's machine ("this e2e test just failed locally", "debug this e2e failure"). From CI it surfaces the test's distinct failure modes from history and pulls evidence for one; locally it reads the run's own trace, snapshot, and logs. Either way it reasons to a falsifiable root cause with the engineer and lands on a test fix, a product-bug issue, or an accepted-flake note. Do NOT use for a test you are writing or currently editing (author-e2e-tests); a whole CI run's failures, or a run ID/URL (e2e-failure-analyzer); Vitest or extension-host failures.

77

Quality

96%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

92%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-engineered orchestrator skill: fully executable commands with explicit verdict contracts at every step, a clearly gated phase workflow with checkpointing and error-recovery, and exemplary progressive disclosure with nine real one-level-deep references. The only weakness is minor redundancy where a few rules and inline mechanisms are restated or belong in the references, keeping conciseness just below the top anchor.

DimensionReasoningScore

Conciseness

The body is dense rules and verbatim commands with no explanations of concepts Claude already knows, but a few rules are restated across sections (pattern-selection appears in Non-negotiable rules, step 7, and the investigate section; the "show the evidence block" rule appears twice) and some inline mechanisms (the merged-fix denominator logic, Last-seen rendering like `5d ago (Jul 24)`) exceed the file's own "invariant here, mechanism in the reference" split. These are minor trimmable instances, matching level 4 rather than the every-token-earns-its-place level 5.

4 / 5

Actionability

Every stage carries a copy-paste-ready command with flags and expected output shapes: `resolve-test-key.js '<whatever they gave>'`, `triage-history.js --test-key ... --lookback-days 14`, `checkpoint.js --triage-id <id> --init --test-key '<key>'`, `record-diagnosis.js --pr <n> --outcome fix-test`. Verdict handling (`stop: true`, `cause: "missing-api-key"`, `no-results`, `open-attempt-in-flight`) is specific, and the exact table format to present is shown. Specific examples cover the common cases.

5 / 5

Workflow Clarity

The multi-step process is explicitly sequenced (resolve identity → history → checkpoint init → prior-triage check → pattern table → selection → evidence → diagnosis → fix approach gate → reproduce → record outcome) with validation throughout: API-key pre-flight, `--read` validation on resume, the checkpoint gate that refuses `phase=done` without a recorded diagnosis, the RED bar for regression tests, and error-recovery paths (candidates → ask which, 403/expired report handling, script-fallbacks reference).

5 / 5

Progressive Disclosure

SKILL.md is a genuine overview holding invariants while nine real reference files (all verified to exist) each own a complete mechanism, every one clearly signaled with what it contains ("owns the verdict table, the run-it offer for no-results, what local evidence cannot answer"). References are one level deep with only lateral sibling links, the scripts table points to references/scripts.md for flags and output contracts, and the file states its own split policy ("this file holds the invariant, the reference holds the mechanism").

5 / 5

Total

19

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An exemplar description: it names the niche precisely, quotes the natural phrases users would actually say, states capabilities for both CI and local entry points, and closes with explicit do-not-use boundaries against named sibling skills. Every clause is a trigger, capability, or boundary with no padding.

DimensionReasoningScore

Specificity

Lists multiple concrete actions covering both entry paths and outcomes: "surfaces the test's distinct failure modes from history and pulls evidence for one", "reads the run's own trace, snapshot, and logs", "reasons to a falsifiable root cause", "lands on a test fix, a product-bug issue, or an accepted-flake note". Coverage is comprehensive across CI, local, and outcome space, so it exceeds the 'minor gaps' level 4.

5 / 5

Completeness

Opens with an explicit "Use when..." clause containing concrete trigger phrases for both CI and local contexts, and explicitly states what it does with multi-clause capability coverage. Both what and when are clearly and explicitly answered.

5 / 5

Trigger Term Quality

Quotes natural user utterances verbatim ("why is <test> failing on main", "this test is flaky in CI", "this e2e test just failed locally", "debug this e2e failure", "is this a test bug or a product bug") alongside synonyms (e2e, Playwright, flaky/flaking, CI, locally). Trigger coverage including quoted phrasings exceeds the 'a few natural terms missing' level 4.

5 / 5

Distinctiveness Conflict Risk

Occupies a clear niche (a specific named Positron e2e/Playwright test failing or flaking) and explicitly disambiguates against named sibling skills: "Do NOT use for a test you are writing or currently editing (author-e2e-tests); a whole CI run's failures, or a run ID/URL (e2e-failure-analyzer); Vitest or extension-host failures." Minimal conflict risk.

5 / 5

Total

20

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
posit-dev/positron
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.