CtrlK
BlogDocsLog inGet started
Tessl Logo

e2e-failure-analyzer

Analyze e2e test failures from a GitHub Actions run. Provide a run ID or URL to download reports, extract traces/screenshots/logs, identify root causes, and get suggested actions. Works with both posit-dev/positron and posit-dev/positron-builds repos.

62

Quality

73%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.claude/skills/e2e-failure-analyzer/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

77%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A strong, execution-grade workflow document: fully executable commands, field-level output contracts, explicit decision rules, and genuine validation checkpoints. Its two weaknesses are the missing `rubric.md` file, which leaves the skill's central root-cause analysis guidance dangling, and some repetition of the screenshot-reading instructions across the two paths.

Suggestions

Ship `rubric.md` alongside SKILL.md (or inline the root-cause categories and evidence-reading order) — the body cites it eight times as the single source of truth but the file is absent from the bundle, so local runs cannot follow Step 7 as written.

State the screenshot-reading and error-context-first rules once in a shared section and reference them from Path A and Path B instead of repeating them near-verbatim per path.

Consider moving the DOM-presence/console-digest interpretation detail into a reference file (or the rubric) so SKILL.md reads as the overview, keeping only the decision rule inline.

DimensionReasoningScore

Conciseness

Nearly every line is task-specific knowledge Claude would not have (output JSON field semantics, decision rules like "an `[after deadline]` line cannot be the cause of the failure", "A `skipped` conclusion means that shard's tests never ran"), and no basic concepts are re-taught. Minor trimming is possible: the "Read all screenshots in a single message" and error-context-first instructions each repeat near-verbatim across Path A and Path B, and the `[analysis rubric](rubric.md)` pointer appears eight times.

4 / 5

Actionability

Commands are copy-paste ready with all flags (`e2e-gather-run-info.js <RUN_URL>`, `e2e-process-project.js --download --run-id ... --cleanup`), output fields are documented item-by-item, and fallback invocations, exact log-file paths, and example History lines are given. The one material gap: Step 7 delegates all root-cause categorization to `rubric.md` ("the single source of truth for the root-cause categories"), which is not present in the bundle, so the skill's central analysis guidance cannot actually be read and executed.

4 / 5

Workflow Clarity

The sequence is explicit and gated: Step 1 gather -> project-list branch (non-empty -> Path A, empty -> Path B) -> per-project/per-job processing -> optional history (with "if the API is unavailable, skip it") -> Step 7 analysis -> cleanup with exact paths (no globs). Validation checkpoints are explicit and repeated where they matter ("Read `failedJobs[].steps` before you touch any test evidence", `total_runs: 0` means a key mismatch not a clean test, check `environment_breakdown` before concluding "flaky"), with warning-signaling and fallback paths for error recovery.

5 / 5

Progressive Disclosure

Sectioning is good and `scripts/README.md` is a real, well-signaled one-level-deep reference (verified to cover the auth, test-key, and response-reading topics it is cited for). But the core reference `[analysis rubric](rubric.md)` is dangling — no `rubric.md` exists in the skill directory while the body positions it as the source of truth — so the heaviest reference cannot be navigated to, and substantial analysis decision-rule content (DOM-presence/console-digest interpretation) stays inline in SKILL.md.

3 / 5

Total

16

/

20

Passed

Description

70%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A specific, well-anchored description with concrete actions and a clearly bounded niche. Its main weakness is the absent "Use when..." trigger clause, which caps completeness and slightly weakens natural triggering despite good keyword coverage.

Suggestions

Add an explicit trigger clause, e.g. "Use when a GitHub Actions CI run fails and you need to triage e2e test failures, or when the user mentions a failed run ID/URL."

Include the natural terms users actually say — "CI run", "Playwright", "flaky" — to broaden trigger coverage.

Optionally mention report merging or historical flaky-vs-regression checks to round out the action coverage toward the top anchor.

DimensionReasoningScore

Specificity

"Provide a run ID or URL to download reports, extract traces/screenshots/logs, identify root causes, and get suggested actions" lists several concrete, distinct actions with clear scoping to "e2e test failures from a GitHub Actions run". Not a 5: coverage has minor gaps — merging sharded blob reports, historical/flaky analysis, and non-e2e job failures are all in the body but invisible in the description.

4 / 5

Completeness

The "what" is clear and concrete, but there is no "Use when..." clause or equivalent explicit trigger guidance — the "when" is only weakly implied by "Provide a run ID or URL". Per the judging guidelines a missing explicit trigger clause caps completeness at 3; it is not a 4 because nothing tells a user (or Claude) when to reach for this skill versus another.

3 / 5

Trigger Term Quality

Natural phrases a user would say are present: "e2e test failures", "GitHub Actions run", "run ID or URL", "root causes". A few common terms users naturally use are missing — "CI" (the everyday word for a GitHub Actions run), "Playwright", and "flaky" — which keeps it below the comprehensive-synonym anchor.

4 / 5

Distinctiveness Conflict Risk

The niche is sharp and repo-pinned: e2e test failure analysis for GitHub Actions runs in "both posit-dev/positron and posit-dev/positron-builds repos". These triggers are unlikely to fire for any other skill, satisfying the clear-niche/minimal-conflict anchor; it does not fall to 4 since no closely related overlapping skill territory is implicated by the description itself.

5 / 5

Total

16

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 6 missing

Warning

Total

15

/

16

Passed

Repository
posit-dev/positron
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.