CtrlK
BlogDocsLog inGet started
Tessl Logo

ui-verify

Makes a UI pull request autonomously verifiable. `author` writes a Markdown intent spec — the happy path plus out-of-bounds cases from a PR-scoped brainstorm — into the PR description as a collapsed, machine-findable block (delegated to by `create-pr` on UI diffs). `run` resolves the PR's live preview URL via the GitHub deployments API and runs the spec there through `aw-tester`, which may adapt the route but checks every expected outcome with evidence, and reports a verdict plus full-page screenshots (`--no-screenshots` opts out). It then tries to break the same change — hostile input, double submits, failed requests, small viewports — and documents every probe with screenshots, never changing the verdict (`--no-adversarial` opts out). `verify` is author-if-needed, then run. `setup` delegates to `aw-setup --target preview`. A LoreKit loop feeds runner friction into the next spec. Web only. Triggers on "write a preview spec", "verify this PR's preview", "try to break this PR's preview", "/ui-verify".

66

Quality

83%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

73%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is exceptionally actionable with well-sequenced operations and strong validation checkpoints, but it fails structurally: it presents itself as a thin index while the rules/ and templates/ files it links to throughout are missing from the bundle, leaving the actual procedures unreachable. Secondary redundancy around Agent0 handling and repeated inconclusive strings could be tightened. Fixing the broken reference targets is the dominant issue.

Suggestions

Ship the referenced bundle files: create rules/ (spec-format.md, runner.md, preview-url-resolution.md, adversarial.md, out-of-bounds.md, memory.md, spec-sources.md, agent0-runtime.md, preview-auth.md) and templates/ so the 'thin index' links resolve, or inline the missing procedures and stop claiming the detail lives elsewhere.

Consolidate the Agent0 unprepared-host explanation into one place (e.g. rules/agent0-runtime.md) and reference it from run/verify/author instead of restating the scenario and the 'not authored (Agent0 sandbox not prepared: …)' line three times.

State the 'inconclusive: no access path for deployment lookup (pass --url)' rule once (Step 0 or rules/preview-url-resolution.md) and reference it from the run outline and verify, rather than repeating the full string at each mention.

DimensionReasoningScore

Conciseness

The body is dense and assumes Claude's competence (no basic-concept explanations), but for a file that declares 'This SKILL.md is a thin index', it carries noticeable repetition: the Agent0 unprepared-host scenario is explained three times (the 'An unprepared Agent0 host' paragraph, the 'author stops with not authored' passage, and its restatement inside 'verify'), and 'inconclusive: no access path for deployment lookup (pass --url)' appears in Step 0, the run outline, and verify. This fits anchor 3 ('mostly efficient but includes some unnecessary explanation or could be tightened') better than anchor 4, whose 'minor instances' understates the redundancy; it is above anchor 2 because most length is genuinely load-bearing operational detail, not padding.

3 / 5

Actionability

Guidance is fully executable: the exact gate command 'node ${CLAUDE_SKILL_DIR}/scripts/is-ui-diff.mjs --base "$(git merge-base origin/HEAD HEAD)"' (script present in scripts/), an mcp-equivalents table ('gh pr view <pr> --json body' → 'mcp__github__pull_request_read with method: "get"'), the complete 80-line on-demand-install bash block with sentinel and budget, exact marker strings ('<!-- ui-verify:v2 -->'), exact report strings, and a decision table keyed on the script's last output line. Specific examples cover the common cases — anchor 5.

5 / 5

Workflow Clarity

Each operation is a numbered sequence with explicit validation checkpoints and feedback loops: author's Step 0 mechanical is-ui-diff gate ('UI_DIFF: no → stop and report not authored'), terminal 'inconclusive' outcomes that must stop without a verdict, the install-sentinel that prevents a second multi-minute attempt, idempotency of verify ('a second verify on the same PR reuses the block'), and hard rules for failure handling ('A missing sandbox browser is NOT RUN (…) … never red'). This matches anchor 5's explicit validation, error-recovery loops, and checklists.

5 / 5

Progressive Disclosure

The body repeatedly delegates its core procedures to files that are not in the bundle: 'Full procedure: rules/runner.md' plus a dozen links to ./rules/*.md (spec-format, preview-url-resolution, runner, adversarial, out-of-bounds, memory, spec-sources, agent0-runtime, preview-auth) and ./templates/*.md — yet the bundle contains only references/ (3 files) and scripts/ (1 file); no rules/ or templates/ directory exists, and even references/adversarial-testing.md opens by pointing to the missing ../rules/adversarial.md. Per the guideline to score against the actual bundle structure, the navigation is broken for the majority of the skill's detail, which is the anchor-2 failure mode (references unusable, needed content unreachable) rather than anchor 3's 'references present but not clearly signaled' — the signals are clear, the targets are absent.

2 / 5

Total

15

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is highly specific, third-person, and explicitly pairs comprehensive 'what' coverage with concrete trigger phrases and scope limits. Its only weaknesses are moderate length and a few missing natural synonyms ('test the preview', 'check the UI') in the trigger list. It is decisively a strong description for autonomous selection.

DimensionReasoningScore

Specificity

The description lists many concrete, distinct actions: 'writes a Markdown intent spec — the happy path plus out-of-bounds cases … into the PR description as a collapsed, machine-findable block', 'resolves the PR's live preview URL via the GitHub deployments API and runs the spec there through aw-tester', 'It then tries to break the same change — hostile input, double submits, failed requests, small viewports'. Coverage is comprehensive across all four operations with opt-out flags named. Not below 4 because there are no meaningful gaps in capability coverage; anchor 5's 'multiple specific concrete actions; comprehensive coverage' matches exactly.

5 / 5

Completeness

Both 'what' and 'when' are answered explicitly: the four operations (author, run, verify, setup) are each described concretely, and the when is stated directly via 'Triggers on …' plus scope ('Web only', 'delegated to by create-pr on UI diffs'). This matches anchor 5's 'clearly and explicitly answers both what AND when with concrete trigger phrases'.

5 / 5

Trigger Term Quality

Trigger phrases are natural and skill-specific: 'Triggers on "write a preview spec", "verify this PR's preview", "try to break this PR's preview", "/ui-verify"' — phrases a user would plausibly say. It falls short of anchor 5 because common variations like 'test the preview', 'check the UI', or 'preview deployment' are absent, leaving a few natural synonyms uncovered. Not 3, because the included terms are genuinely natural rather than technical jargon.

4 / 5

Distinctiveness Conflict Risk

The niche is unambiguous — attaching machine-findable verification specs to UI pull requests and running them against live preview deployments — a combination no generic skill would claim, and the trigger phrases are unique to it ('verify this PR's preview'). Minimal conflict risk, matching anchor 5. The minor caveat that the ~200-word length dilutes scannability does not create overlap with other skills.

5 / 5

Total

19

/

20

Passed

Validation

68%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 11 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

relative_links

Relative link issues: 39 missing, 10 suspicious

Warning

referenced_paths_exist

Referenced path issues: 3 missing

Warning

Total

11

/

16

Passed

Repository
mthines/agent-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.