CtrlK
BlogDocsLog inGet started
Tessl Logo

webapp-testing

Toolkit for interacting with and testing local web applications using Playwright. Supports verifying frontend functionality, debugging UI behavior, capturing browser screenshots, and viewing browser logs.

91

2.17x
Quality

72%

Does it follow best practices?

Impact

100%

2.17x

Average score across 7 eval scenarios

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./app/skills/webapp-testing/SKILL.md

The canonical home for this skill is webapp-testing in anthropics/skills

SKILL.md
Quality
Evals
Security

Quality

Content

78%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A strong, highly actionable body: executable commands and code, a clear decision tree, and good black-box handling of the bundled script. The main defects are the missing examples/ files referenced in the Reference Files section (broken navigation), the absence of explicit post-action validation steps, and some repetition of the networkidle advice.

Suggestions

Fix the broken references: either add the three promised files under examples/ (element_discovery.py, static_html_automation.py, console_logging.py) or remove the 'Reference Files' section — pointing at nonexistent files defeats navigation.

Add explicit validation checkpoints to the reconnaissance-then-action pattern (e.g., verify an expected element/text appears after an action, and how to recover if a discovered selector fails).

Consolidate the networkidle guidance — it is currently stated in the decision tree, a code comment, the Common Pitfall section, and Best Practices; one authoritative statement plus the pitfall callout would tighten the body.

DimensionReasoningScore

Conciseness

The body is dense with actionable material and assumes Claude's competence — no time is spent explaining what Playwright or web testing is. Minor trimming is possible: the networkidle guidance is repeated in the decision tree, a code comment, the Common Pitfall section, and Best Practices, and the black-box-scripts advice appears in both the intro and Best Practices; a typo ('abslutely') also remains. This sits at anchor 4 (efficient with minor over-explanation) rather than 5, and well above anchor 3's 'noticeably could be tightened'.

4 / 5

Actionability

Copy-paste-ready bash commands cover both single- and multi-server cases, and the Playwright snippet plus reconnaissance code ('page.screenshot(...)', "page.locator('button').all()", 'wait_for_load_state') are fully executable and match the real bundled script's interface. The common cases are covered concretely, matching the anchor-5 example.

5 / 5

Workflow Clarity

The decision tree sequences static-vs-dynamic and server-running-or-not branches into a concrete reconnaissance-then-action pattern (navigate → screenshot/inspect → identify selectors → execute), which is clear and well-ordered. It falls short of anchor 5 because there are no explicit validation checkpoints or feedback loops (e.g., verifying an expected element or state after an action, or what to do when a selector is not found), though the non-destructive nature of the task means the destructive-operation cap does not apply.

4 / 5

Progressive Disclosure

The body itself is well structured and the main bundle reference (scripts/with_server.py) is real, correctly signposted, and used as a black box. However, the 'Reference Files' section advertises an examples/ directory with three named files (element_discovery.py, static_html_automation.py, console_logging.py) that do not exist in the bundle, so a core navigation pointer is broken — matching anchor 3 (references present but unreliable, structure could be better) rather than anchor 4.

3 / 5

Total

16

/

20

Passed

Description

66%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A solid, third-person description with a clear 'what' and good natural keyword coverage, held back mainly by the complete absence of a 'when to use' clause. Adding an explicit trigger sentence (e.g., 'Use when testing locally running web apps, taking browser screenshots, or debugging frontend behavior') would lift completeness and distinctiveness.

Suggestions

Add an explicit 'Use when...' clause naming concrete trigger situations (e.g., 'Use when the user wants to test a locally running web app, capture browser screenshots, inspect console logs, or debug frontend behavior').

Include common synonyms users actually say — 'browser automation', 'end-to-end (e2e) testing', 'UI testing' — to broaden trigger-term coverage.

Make 'verifying frontend functionality' and 'debugging UI behavior' more concrete (e.g., 'click flows and form submissions', 'inspecting rendered DOM state') to reach full specificity.

DimensionReasoningScore

Specificity

The description names the domain ('local web applications using Playwright') and lists four capability areas ('verifying frontend functionality, debugging UI behavior, capturing browser screenshots, and viewing browser logs'), two of which are highly concrete. It falls just short of the comprehensive anchor 5 because 'verifying frontend functionality' and 'debugging UI behavior' are generic compared to anchor-5 examples that enumerate fully concrete actions.

4 / 5

Completeness

The 'what' is clearly stated (a Playwright toolkit with four listed capabilities), but there is no 'Use when...' clause or equivalent trigger guidance anywhere in the description, which caps completeness at 3 per the judging guidelines. It is above anchor 2 because the 'what' is explicit and multi-part rather than vague.

3 / 5

Trigger Term Quality

Natural keywords like 'web applications', 'testing', 'Playwright', 'screenshots', and 'browser logs' match what users would say, giving good coverage. It is not anchor 5 because common variations such as 'browser automation', 'end-to-end/e2e testing', 'UI testing', and 'localhost' are absent.

4 / 5

Distinctiveness Conflict Risk

Scoping to 'local web applications' plus the Playwright name carves a mostly distinct niche with clear triggers, matching anchor 4. It is not anchor 5 because it could still overlap with sibling browser-automation or general testing skills given the absence of explicit 'use when' boundary phrasing.

4 / 5

Total

15

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
ZHangZHengEric/Sage
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.