CtrlK
BlogDocsLog inGet started
Tessl Logo

aw-tester-chrome

Runs UI verification specs in-session through the claude-in-chrome extension — the Chrome sibling of the aw-tester agent. Reads the same specs.md (WHEN/THEN grammar or a Markdown intent spec) and aw-target.yml, drives Chrome interactively (navigate → read → act → assert, seeing the page between steps), and emits the same compact verdict block. Runs in the current session against an already-logged-in Chrome, so it is faster than the Playwright sub-agent for local runs, but it needs the browser extension and does not work in remote / CI envs. Falls back to aw-tester when the extension is not connected. Triggers on "run the spec in chrome", "verify with the chrome driver", "aw-tester-chrome", or a "ui-verify run --driver chrome" dispatch.

72

Quality

90%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An unusually well-engineered skill body: fully executable commands, explicit validation checkpoints and fallbacks, and disciplined delegation of shared rules to an external contract. The only weaknesses are modest local restatement of contract rules (token cost) and reliance on referenced files that are not present in the bundle.

DimensionReasoningScore

Conciseness

The body is dense and operational with essentially no explanation of concepts Claude already knows — every section carries instructions, tool names, or verdict rules. It is not a 5 because at ~300 lines it restates some contract-owned rules locally (e.g., route replay/heal semantics in the Intent specs section) that could be trimmed or left entirely to the referenced contract.

4 / 5

Actionability

Copy-paste-ready guidance throughout: a concrete ToolSearch 'select:...' string, memory.list calls with scopes and tags, an exact preflight fallback YAML block, a per-assertion tool mapping (read_page, get_page_text, read_network_requests), exact paths and caps (AUTO_CAPTURE_CAP = 30), and a defined verdict schema. Fully executable; anchor 5.

5 / 5

Workflow Clarity

A clearly sequenced multi-step process — load tools, preflight the extension, read cross-run lessons, parse inputs, auth check, execution loop, verdict, lessons — with explicit validation checkpoints (preflight probe, login-screen detection, healing retry ladder, bail mode) and honest inconclusive/fallback failure handling. Anchor 5 with feedback loops.

5 / 5

Progressive Disclosure

Good structure: shared rules are clearly delegated to a well-signaled, one-level-deep contract ('Read the spec-run contract first. It owns the locator ladder, the auth semantics...'), and each section states what it owns vs. defers. However, the bundle contains no reference files at all — every referenced path (../rules/spec-run-contract.md, ../templates/aw-tester.agent.md, ../../../quality/jev-assert/SKILL.md) points outside the skill directory and cannot be verified from the bundle, leaving navigation dependent on an external repo layout. Minor organization gap; anchor 4.

4 / 5

Total

18

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: concrete third-person capabilities, explicit trigger phrases, and clear boundary/differentiation from the Playwright sibling. Trigger-term coverage is good but leans on internal ecosystem dispatch strings rather than a full spread of natural synonyms.

DimensionReasoningScore

Specificity

The description lists multiple concrete, third-person actions — reads specs.md (WHEN/THEN grammar or intent spec) and aw-target.yml, drives Chrome interactively (navigate → read → act → assert), emits a compact verdict block, and falls back to aw-tester when the extension is not connected. Comprehensive coverage of capabilities; no vague language to place it below the anchor-5 example.

5 / 5

Completeness

It explicitly answers both 'what' (reads specs and target files, drives Chrome interactively, emits a verdict block, falls back to aw-tester) and 'when' via a 'Triggers on ...' clause with concrete trigger phrases. This matches the anchor-5 example pattern; it is not anchor 4 because the when-guidance is fully explicit rather than improvable.

5 / 5

Trigger Term Quality

Four explicit trigger phrases are given, including natural user phrasings ('run the spec in chrome', 'verify with the chrome driver'), but two are internal dispatch strings ('aw-tester-chrome', 'ui-verify run --driver chrome') and common variations/synonyms such as 'test in chrome' or 'browser verification' are absent. Good coverage with a few natural terms missing — anchor 4, not the comprehensive synonym coverage of anchor 5.

4 / 5

Distinctiveness Conflict Risk

It carves a clear niche — the Chrome sibling of the aw-tester Playwright sub-agent — and explicitly states the boundary (needs the claude-in-chrome extension, does not work in remote/CI envs, falls back to aw-tester when disconnected). Minimal conflict risk with the sibling skill; anchor 5.

5 / 5

Total

19

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

relative_links

Relative link issues: 14 suspicious

Warning

Total

13

/

16

Passed

Repository
mthines/agent-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.