CtrlK
BlogDocsLog inGet started
Tessl Logo

webapp-testing

Toolkit for interacting with and testing local web applications using Playwright. Supports verifying frontend functionality, debugging UI behavior, capturing browser screenshots, and viewing browser logs.

91

2.17x
Quality

72%

Does it follow best practices?

Impact

100%

2.17x

Average score across 7 eval scenarios

SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./skills/anthropic-webapp-testing/SKILL.md

The canonical home for this skill is webapp-testing in anthropics/skills

SKILL.md
Quality
Evals
Security

Quality

Content

78%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is highly actionable with verified, copy-paste-ready commands and a clear decision tree for choosing an approach. Its weaknesses are a referenced examples/ directory missing from the bundle and minor redundancy in the black-box script guidance.

Suggestions

Ship the examples/ directory (element_discovery.py, static_html_automation.py, console_logging.py) or remove the 'Reference Files' section pointing to it, since the paths currently resolve to nothing.

Consolidate the duplicated black-box-scripts guidance ('DO NOT read the source...' paragraph and the 'Use bundled scripts as black boxes' bullet) into one place, and fix the 'abslutely' typo.

Add an explicit verification step to the reconnaissance pattern (e.g., assert the page title or a key element after networkidle before acting) to close the workflow's validation gap.

DimensionReasoningScore

Conciseness

The body is efficient — a compact decision tree, ready-to-run commands, and one runnable snippet — but the black-box-scripts instruction is repeated ('DO NOT read the source...' and again under 'Best Practices'), and a typo ('abslutely') adds minor trimmable noise.

4 / 5

Actionability

Both bash invocations are copy-paste ready and verified against the actual with_server.py argparse interface (--server/--port/trailing command), and the Python Playwright snippet is complete and executable, covering the common single- and multi-server cases.

5 / 5

Workflow Clarity

The decision tree sequences static-vs-dynamic and running-vs-not-running paths with a failure branch ('Fails/Incomplete → Treat as dynamic'), and reconnaissance-then-action is a numbered 4-step sequence, but validation checkpoints are implicit (wait for networkidle) rather than explicit verify/retry steps.

4 / 5

Progressive Disclosure

Sections are well-organized and scripts/with_server.py is a real, working reference, but the 'Reference Files' section points to an examples/ directory with three named files (element_discovery.py, static_html_automation.py, console_logging.py) that does not exist in the bundle, breaking navigation.

3 / 5

Total

16

/

20

Passed

Description

66%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description names a clear domain and tool with four enumerated capabilities, making it specific and largely distinct. Its main weakness is the absence of any 'Use when...' trigger clause, which caps completeness and limits how reliably users or Claude would select it.

Suggestions

Append an explicit trigger clause, e.g., 'Use when testing or debugging locally running web apps, taking browser screenshots, or inspecting browser console logs.'

Add natural trigger synonyms such as 'end-to-end testing', 'E2E', 'frontend testing', or 'UI debugging' to broaden keyword coverage.

Make the two abstract capabilities more concrete, e.g., replace 'verifying frontend functionality' with 'clicking through user flows and asserting element state'.

DimensionReasoningScore

Specificity

Lists several specific actions ('verifying frontend functionality, debugging UI behavior, capturing browser screenshots, and viewing browser logs') and names the tool (Playwright), but two of the four actions are more abstract than the fully concrete anchor-5 example, leaving minor gaps in coverage.

4 / 5

Completeness

The 'what' is clearly and concretely stated, but there is no 'Use when...' clause or equivalent explicit trigger guidance, which caps completeness at 3 per the judging guidelines.

3 / 5

Trigger Term Quality

Contains natural user terms like 'web applications', 'Playwright', 'browser', 'screenshots', 'logs', and 'UI', giving good keyword coverage; a few natural variations (e.g., 'end-to-end', 'E2E', 'frontend testing') are missing, so it falls short of comprehensive.

4 / 5

Distinctiveness Conflict Risk

'using Playwright' and 'local web applications' carve a mostly distinct niche with minimal conflict risk against unrelated skills, though minor overlap remains with generic browser-automation skills and no explicit trigger phrases separate them.

4 / 5

Total

15

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
boisenoise/skills-collections
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.