CtrlK
BlogDocsLog inGet started
Tessl Logo

e2e-failure-analyzer

Analyze e2e test failures from a GitHub Actions run. Provide a run ID or URL to download reports, extract traces/screenshots/logs, identify root causes, and get suggested actions. Works with both posit-dev/positron and posit-dev/positron-builds repos.

61

Quality

72%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./.claude/skills/e2e-failure-analyzer/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

77%

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is highly actionable with concrete commands, clear output schemas, and a well-sequenced branching workflow. Its weaknesses are moderate verbosity with some duplicated guidance and a progressive-disclosure gap where the referenced rubric.md is absent from the bundle.

Suggestions

Restore or include the referenced rubric.md (or repoint the links to an existing file) — it is cited repeatedly as the 'single source of truth' but is missing from the bundle.

De-duplicate the screenshot/error-context reading guidance between Path A and Path B by stating it once and referencing it from both paths.

Consider moving the large JSON field-shape and DOM-presence/console-digest detail into a reference file so SKILL.md stays a lean overview.

DimensionReasoningScore

Conciseness

Mostly efficient and largely domain-specific (JSON field shapes, REPORT_DIR resolution, DOM-presence ambiguity), but the screenshot-reading instructions are near-duplicated across Path A and Path B and some explanatory passages could be tightened, fitting the 'mostly efficient but could be tightened' anchor rather than the lean 3.

2 / 3

Actionability

Provides fully executable, copy-paste-ready bash commands with exact flags and precise output JSON field descriptions for every step, matching the 'fully executable code/commands; copy-paste ready' anchor.

3 / 3

Workflow Clarity

The multi-step process is clearly sequenced with an explicit decision branch (non-empty projects -> Path A, empty -> Path B), per-project/per-job iteration guidance, fallback options, and a safety-noted cleanup, satisfying the clear-sequence anchor.

3 / 3

Progressive Disclosure

Scripts are clearly listed and one-level-deep references to scripts/README.md are well-signaled, but the heavily-referenced [analysis rubric](rubric.md) is missing from the bundle and a fair amount of detail (DOM-presence/console-digest explanation, full field schemas) is inline rather than split out, fitting the 'some structure but could be better organized' anchor.

2 / 3

Total

10

/

12

Passed

Description

67%

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific and distinctive, naming concrete actions and a well-scoped niche. Its main weakness is the absence of an explicit 'Use when...' trigger clause and incomplete coverage of common trigger terms like 'CI' or 'flaky'.

Suggestions

Add an explicit 'Use when...' clause, e.g. 'Use when a CI/e2e run has failed and you need to triage why.'

Include common trigger variations users would actually say — 'CI failure', 'Playwright failures', 'flaky e2e tests' — to improve trigger term coverage.

Optionally note when NOT to use it (e.g. deep single-test flake investigation belongs to triage-e2e-test) to further reduce conflict risk.

DimensionReasoningScore

Specificity

Lists multiple concrete actions — 'Provide a run ID or URL to download reports, extract traces/screenshots/logs, identify root causes, and get suggested actions' — matching the anchor for enumerating several specific actions.

3 / 3

Completeness

Clearly answers 'what' (download/extract/identify/suggest) but the 'when' is only implied — there is no 'Use when...' clause or equivalent explicit trigger guidance, which per the judging guidelines caps completeness at 2.

2 / 3

Trigger Term Quality

Includes relevant natural terms like 'e2e test failures' and 'GitHub Actions run', but misses common variations a user would say such as 'CI', 'Playwright', or 'flaky tests', fitting the 'some relevant keywords but missing common variations' anchor.

2 / 3

Distinctiveness Conflict Risk

A clear niche scoped to 'e2e test failures from a GitHub Actions run' for two named repos (posit-dev/positron and posit-dev/positron-builds), giving distinct triggers unlikely to conflict with other skills.

3 / 3

Total

10

/

12

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 5 missing

Warning

Total

15

/

16

Passed

Repository
posit-dev/positron
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.