CtrlK
BlogDocsLog inGet started
Tessl Logo

agui-playwright-validate

Manually validate a running web app or docs site with the Playwright MCP server — navigate pages, assert headings/text/status, check for console errors, hover/click to exercise interactions, take screenshots for human review, and verify visual details such as an icon's color in light vs dark mode by reading computed styles. USE FOR: "validate with Playwright", "open the browser and check", visually verifying a UI or docs change, confirming a CSS/SVG/icon edit rendered, checking a page for console errors, comparing light/dark mode appearance. DO NOT USE FOR: writing automated unit/E2E test files (dojo e2e lives in agui-dojo); docs-site content/structure checks specific to the .NET SDK docs (use agui-dotnet-sdk-docs). INVOKES: browser_navigate, browser_snapshot, browser_take_screenshot, browser_evaluate, browser_hover, browser_click, browser_console_messages Playwright MCP tools.

76

Quality

94%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A strong, experience-derived skill: an ordered validation loop with real checkpoints and a set of gotchas (cache, overlays, ephemeral ids, theme-relative color assertions) an agent would otherwise rediscover slowly. The only meaningful improvement is removing the duplication between the inline ❌ callouts and the Anti-patterns section.

Suggestions

Remove the Anti-patterns entries that repeat inline ❌ callouts verbatim (soft-reload trust, script-wrapping, artifact/dev-server leftovers), keeping that section only for pitfalls not already flagged in place.

Consolidate cleanup guidance into the Cleanup discipline section and reference it from the loop, rather than restating it as anti-patterns.

Consider a one-line "Evidence to report" checklist after step 5 of the loop to make the reporting requirement scannable without reading the full technique sections.

DimensionReasoningScore

Conciseness

The body is dense with non-obvious gotchas (ephemeral evaluate-injected ids, cache-busted reloads, hover target requirements) and avoids explaining known concepts, but the ❌ callouts are duplicated — the soft-reload warning and cleanup/artifact warnings each appear both inline in their technique sections and again in the Anti-patterns section — which is trimmable without losing anything.

4 / 5

Actionability

Guidance is fully executable: exact MCP tool names (browser_navigate, browser_snapshot, browser_console_messages, browser_press_key Escape), copy-paste browser_evaluate JS for computed-style comparison and theme toggling, and a concrete cache-busting pattern (?v=<timestamp>) — no pseudocode or vague direction.

5 / 5

Workflow Clarity

The validation loop is a clearly sequenced cheap-to-expensive procedure with "Stop as soon as you have a confident answer", per-step assertions ("assert 0 errors ... Do this on every page you validate"), error-recovery feedback loops (Escape on intercepted clicks, re-query from a fresh snapshot after re-render), and evidence-based pass/fail reporting; no destructive or batch operations that would cap the score.

5 / 5

Progressive Disclosure

Sections are well organized with clear headers and nothing that clearly belongs in a separate file is inlined (no bundle files exist, appropriately). It falls short of the 5 anchor — which rewards content appropriately split across well-signaled references and easy navigation — because the skill is ~130 lines and the duplicated inline ❌ callouts vs. the Anti-patterns section slightly muddy navigation.

4 / 5

Total

18

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An exemplary description: concrete capability list, explicit natural-language triggers, and explicit negative boundaries that separate it from sibling skills. Its only cost is length (~150 words), but every clause carries information rather than padding.

DimensionReasoningScore

Specificity

The description lists six concrete actions covering the full validation workflow — "navigate pages, assert headings/text/status, check for console errors, hover/click to exercise interactions, take screenshots for human review, and verify visual details such as an icon's color in light vs dark mode by reading computed styles" — matching the comprehensive-coverage anchor, not the minor-gaps (4) anchor.

5 / 5

Completeness

It explicitly answers both what (the first sentence's capability list) and when ("USE FOR: ..." with concrete trigger phrases), mirroring the anchor-5 example structure exactly; the 4 anchor requires a 'when' that could be more explicit, which does not apply here.

5 / 5

Trigger Term Quality

The USE FOR list gives natural phrases users would actually say ("validate with Playwright", "open the browser and check", "confirming a CSS/SVG/icon edit rendered", "comparing light/dark mode appearance") with synonym coverage; it fits the comprehensive natural-terms anchor rather than the few-missing (4) anchor.

5 / 5

Distinctiveness Conflict Risk

A clear niche (manual Playwright-MCP validation) is reinforced by active disambiguation — "DO NOT USE FOR: writing automated unit/E2E test files (dojo e2e lives in agui-dojo); docs-site content/structure checks specific to the .NET SDK docs (use agui-dotnet-sdk-docs)" — so conflict risk with sibling skills is minimal.

5 / 5

Total

20

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
ag-ui-protocol/ag-ui
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.