CtrlK
BlogDocsLog inGet started
Tessl Logo

playwright-trace-analyzer

Analyzes Playwright E2E `trace.zip` archives (and bare trace JSONL when unpacked). Extracts the action timeline, network waterfall, console errors, and DOM-snapshot anchors, then identifies the highest-impact problems (flaky waits, slow selectors, network bottlenecks, hung actions, unhandled console errors, navigation churn) and proposes concrete test or app fixes ranked by measured impact. Auto-detects whether the input is a `trace.zip`, a directory of unpacked trace files, or a single `trace.trace` / `trace.network` JSONL stream. Iterates via the `/confidence` skill — if root-cause certainty is below 90%, it digs deeper before recommending a fix. Use when handed a Playwright trace, asked "why is this test flaky?", "why did the test time out?", or asked to optimise an E2E suite with evidence. Triggers on "analyze trace", "playwright trace", "e2e trace", "test flake", "why did playwright fail", "playwright timing", "/playwright-trace-analyzer".

74

Quality

93%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A strong orchestrator-style body: fully executable commands for every input shape, a six-phase gated workflow with a threshold-based confidence feedback loop and a Definition of Done checklist, and disciplined deferral of detail to bundle files. The two weaknesses are structural: the phase-to-files mapping is duplicated across two tables (with the confidence policy stated three times), and 8 of the 14 referenced bundle paths (all of `rules/` and `templates/`) are absent from the provided bundle, leaving the core per-phase playbooks unverifiable.

Suggestions

Merge the 'Workflow' phase table and the 'Required reading by phase' table into a single Phase → files → gate table; the two currently duplicate the same rule-file listings per phase and can drift apart.

Ship (or restore) the referenced `rules/*.md` and `templates/analysis-report.md` files in the bundle — 8 of the 14 paths cited in the body resolve to nothing, including the report template required by Phase 5's gate and the Definition of Done.

State the confidence-gate policy once in detail (the 'Confidence-gated iteration' section) and merely cross-reference it from the workflow table and Core principle 5 to remove the triple restatement.

DimensionReasoningScore

Conciseness

The body is lean per-sentence — no explanations of concepts Claude already knows (no 'what is Playwright' padding), and details are delegated to bundle files ("Detailed extraction rules, analysis playbooks, and report templates live under `rules/`, `references/`, and `templates/`. Load only what the current phase needs"). However, there is structural redundancy: the "Workflow" phase table's "Rule file" column and the "Required reading by phase" table list the same rule files per phase, and the confidence-gate policy is stated three times (workflow table row 4, the "Confidence-gated iteration" section, and Core principle 5 "Confidence-gated honesty"). This fits the 4 anchor ('Efficient; minor instances of over-explanation that could be trimmed') rather than 5 ('every token earns its place') because merging the two phase tables and de-duplicating the gate policy would remove ~20 lines without information loss; not 3 because the excess is duplicated structure, not unnecessary explanation.

4 / 5

Actionability

Fully executable, copy-paste-ready commands with placeholders and flags: "node <skill_dir>/scripts/trace-extract.mjs <path/to/trace.zip> [--out <dir>]", "node <skill_dir>/scripts/trace-summary.mjs <dir>", "node <skill_dir>/scripts/trace-diff.mjs <pass-dir> <fail-dir>", "node <skill_dir>/scripts/fetch-gh-run.mjs https://github.com/<owner>/<repo>/actions/runs/<id> [--out <dir>]". Concrete detection signals are given ("Magic bytes `50 4b 03 04`", "entries with `type: 'resource-snapshot'`"), plus a model of what a real finding looks like ("'`page.click('text=Save')` waited 4,820ms across 3 attempts'") and a source-mapping mechanism ("action callId → test file/line (Playwright trace events embed `location: { file, line, column }`)"). This matches the 5 anchor; common input cases (zip, dir, JSONL, GitHub Actions URL) are each covered with a specific command, and the `gh`-missing fallback is handled.

5 / 5

Workflow Clarity

A six-phase workflow with an explicit gate per phase and "Do not skip a gate" (e.g. Phase 0 gate: "Format detected, archive unpacked, `trace.trace` + `trace.network` parseable"; Phase 3 gate: "Each hotspot mapped to a code-level cause... with file path or line where possible"), a threshold-based feedback loop ("≥ 90% → Proceed to Phase 5... 70–89% → Run one deeper pass... < 70% → Surface the gap to the user with a question — do not propose changes on speculation"), an iteration cap ("After two deep-dive iterations without reaching 90%, stop and present findings as a hypothesis"), and a "Definition of Done" checklist for final verification. This matches the 5 anchor (clear sequence, explicit validation steps, feedback loops, checklist). The operations are read-only analysis, so the destructive/batch validation cap does not apply.

5 / 5

Progressive Disclosure

The architecture matches the 5 anchor on design: a self-described "thin orchestrator" body with well-signaled, one-level-deep markdown links organized by phase ("Load on demand — do not preload"), and all six files in the provided bundle's `references/` and `scripts/` directories cited by the body exist (flake-patterns.md, performance-patterns.md, trace-extract.mjs, trace-summary.mjs, trace-diff.mjs, fetch-gh-run.mjs). However, scoring against the actual bundle structure per the judging guideline: 8 of the 14 referenced paths — `rules/input-detection.md`, `rules/measurement-methodology.md`, `rules/action-timing.md`, `rules/network-analysis.md`, `rules/console-and-errors.md`, `rules/flake-diagnosis.md`, `rules/confidence-loop.md`, and `templates/analysis-report.md` — are not present in the bundle (no `rules/` or `templates/` directories exist). Since every phase's method and the final report template live in those files, navigation fails for the most load-bearing references, which drops this to the 4 anchor ('good structure; most content appropriately placed; references mostly clear; minor organization gaps') rather than 5 ('easy navigation'). Not 3: the signaling and one-level-deep organization are excellent, and every referenced file that exists in the bundle resolves correctly.

4 / 5

Total

18

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An exemplary description: third-person, dense with concrete capabilities (extraction targets, diagnosis categories, ranked fix proposal, format auto-detection), an explicit 'Use when...' clause with natural question-form triggers, and comprehensive keyword coverage including file extensions and the slash-command name. No fluff or over-claims despite its length — every clause states a capability or a trigger.

DimensionReasoningScore

Specificity

Quotes: "Extracts the action timeline, network waterfall, console errors, and DOM-snapshot anchors", "identifies the highest-impact problems (flaky waits, slow selectors, network bottlenecks, hung actions, unhandled console errors, navigation churn)", "proposes concrete test or app fixes ranked by measured impact", "Auto-detects whether the input is a `trace.zip`, a directory of unpacked trace files, or a single `trace.trace` / `trace.network` JSONL stream". This lists multiple specific concrete actions with comprehensive coverage of the skill's capability surface, matching the 5 anchor ('Extract text and tables from PDF files, fill forms, merge documents, convert pages to images'). Not 4: coverage has no gaps — extraction, diagnosis, fix proposal, and input auto-detection are all enumerated concretely; not below 5 because there is no generic or abstract language anywhere.

5 / 5

Completeness

The 'what' is explicit and detailed ("Analyzes Playwright E2E `trace.zip` archives... Extracts... identifies... proposes concrete test or app fixes ranked by measured impact") and the 'when' is an explicit 'Use when...' clause with concrete trigger phrases ("Use when handed a Playwright trace, asked 'why is this test flaky?'..."). This matches the 5 anchor, which requires both what AND when with concrete triggers. Not 4: the 'when' clause is not merely present but enumerates specific user scenarios rather than a generic 'Use when working with X'. The description is in third person ('Analyzes', 'Extracts'), per the voice guideline.

5 / 5

Trigger Term Quality

Quotes: "Use when handed a Playwright trace, asked 'why is this test flaky?', 'why did the test time out?'", "Triggers on 'analyze trace', 'playwright trace', 'e2e trace', 'test flake', 'why did playwright fail', 'playwright timing', '/playwright-trace-analyzer'". Natural user phrasings, synonyms, exact file formats (trace.zip, trace.trace, trace.network), and the slash-command form are all present, matching the 5 anchor ('PDF files, PDFs, forms, document extraction, .pdf'). Not 4: no common variation is missing — both question-form and keyword-form triggers are covered.

5 / 5

Distinctiveness Conflict Risk

Quotes: "Analyzes Playwright E2E `trace.zip` archives", "playwright trace", "why did playwright fail". The niche (Playwright trace forensics) is distinct and nearly every trigger is Playwright-anchored, matching the 5 anchor's 'clear niche with distinct triggers; minimal conflict risk'. Not 4: the only broadly-worded trigger is "test flake", but it appears alongside Playwright-specific anchors and the file-format signals (trace.zip, trace.network), so overlap with a generic flake skill is minimal.

5 / 5

Total

20

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

relative_links

Relative link issues: 21 missing

Warning

Total

13

/

16

Passed

Repository
mthines/agent-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.