CtrlK
BlogDocsLog inGet started
Tessl Logo

playwright-app-testing

Test the Expensify App using Playwright browser automation. Use when user requests browser testing, after making frontend changes, or when debugging UI issues

65

Quality

82%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

High

Do not use without reviewing

SKILL.md
Quality
Evals
Security

Quality

Content

75%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A tight, highly actionable skill body: concrete commands, credentials, and a clear sequenced workflow with sensible checkpoints. The main improvement opportunities are trimming the redundant Example Usage section and adding one concrete Playwright MCP interaction example.

DimensionReasoningScore

Conciseness

The body is lean — bare commands, no explanation of concepts Claude already knows — but has minor trimming opportunities: the "Example Usage" scenarios largely restate the "When to Use" triggers, and "Dev Server Details" is a heading wrapping a single URL line. Fits 'efficient; minor instances of over-explanation that could be trimmed' rather than the every-token-earns-its-place anchor.

4 / 5

Actionability

Guidance is copy-paste ready throughout: `ps aux | grep "rspack"`, `cd App && npm run web`, the exact dev URL, the random-Gmail sign-in pattern, "Security code: Always `000000`", and exact `sed`/`grep` commands for the SKIP_ONBOARDING flag. The one high-level gap is "Use Playwright MCP tools to inspect, click, type, and navigate", which lacks a concrete usage example — matching 'mostly executable; minor gaps' rather than fully covering common cases.

4 / 5

Workflow Clarity

The workflow is clearly sequenced (verify server → navigate → interact) with a prerequisites check, a verification checkpoint ("take a snapshot to check the result and only add short waits if the page hasn't updated"), and check/revert-plus-restart guidance around the env-flag change. It stops short of anchor 5's explicit validate→fix→retry feedback loops, so 'clear sequence with most checkpoints; minor validation gaps' is the best fit.

4 / 5

Progressive Disclosure

No bundle files exist, so this is scored on structure alone: well-organized sections with clear headers, no wall of text, no buried references, and inline sign-in detail at a size where that is appropriate. At ~75 lines it slightly exceeds the under-50-line simple-skill case, and the sign-in/env-flag detail could move to a reference file if the skill grows, matching 'good structure; minor organization gaps'.

4 / 5

Total

16

/

20

Passed

Description

82%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description that clearly and explicitly states what the skill does and when to use it, with a well-differentiated niche. The only weakness is that the capability statement is compressed to a single action rather than enumerating the specific testing actions the skill covers.

Suggestions

Expand the capability statement to name a few concrete actions, e.g., "Test the Expensify App with Playwright: navigate, inspect elements, click, type, and verify sign-in flows in a browser."

Add common synonym triggers such as "e2e testing", "end-to-end testing", or "automated UI testing" to broaden natural keyword coverage.

DimensionReasoningScore

Specificity

"Test the Expensify App using Playwright browser automation" names the domain plus one concrete action but stops there — it does not list several specific actions (e.g., inspect elements, click, verify sign-in flows), matching the '1-2 concrete actions, not comprehensive' anchor rather than the 'several specific actions' anchor above.

3 / 5

Completeness

Explicitly answers both: what ("Test the Expensify App using Playwright browser automation") and when ("Use when user requests browser testing, after making frontend changes, or when debugging UI issues") with concrete trigger phrases, matching the top anchor exactly.

5 / 5

Trigger Term Quality

Includes natural phrases users would say — "browser testing", "frontend changes", "debugging UI issues", plus "Playwright" — but misses common variations like "e2e testing", "end-to-end", or "automated testing", so it fits 'good keyword coverage; a few natural terms missing' rather than the comprehensive-synonym anchor.

4 / 5

Distinctiveness Conflict Risk

Scoped to a specific app (Expensify) and tool (Playwright), giving it a clear niche with distinct triggers and minimal risk of firing for unrelated skills.

5 / 5

Total

17

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
Expensify/App
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.