CtrlK
BlogDocsLog inGet started
Tessl Logo

aw-setup

One-time (but safely re-runnable) setup flow that scaffolds a project's aw-tester aw-target: detects auth strategy, captures storage state, writes .claude/aw-targets/<target>.yml, and validates with a smoke spec. `--target local` (default) scaffolds the local dev target; `--target preview` scaffolds the PR-preview target ui-verify runs against (and is what `/ui-verify setup` delegates to). Also records the repo's UI surface (what counts as a UI change) so the is-ui-diff gate is accurate per repo, and the preview auth profile. Re-runs detect the existing aw-target and only re-prompt for what broke or changed. Triggers on "/aw-setup", "setup aw-tester", "scaffold aw-target".

64

Quality

81%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

70%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body delivers an exceptionally clear, validated workflow with copy-paste commands, but it is long and keeps template-level reference material inline while pointing at bundle paths that don't exist in this distribution. Conciseness and file organization are the two dimensions holding it back.

Suggestions

Move the 'Auth flow templates' section (~105 lines: env-contract tables, YAML snippets, decision tree) into the template files themselves or a references/ file, keeping only the Phase B strategy-selection table in SKILL.md.

Give the local aw-target template an explicit path (as the preview target does) and verify that referenced paths (./templates/, ../templates/aw-tester.agent.md) actually ship in the skill bundle.

Merge the dry-run and re-run examples into one trimmed example — the re-run example largely reprises the dry-run's phases.

DimensionReasoningScore

Conciseness

The ~610-line body is dense and operational (little concept padding Claude already knows), but noticeably long: the ~105-line 'Auth flow templates' section, two full worked examples (dry-run and re-run), a 'Which one wins' decision tree, and 'Compatibility notes' could all be tightened or offloaded. This fits 'mostly efficient but... could be tightened' rather than the 'minor instances' of a 4.

3 / 5

Actionability

Mostly executable guidance: exact probe commands ("AUTH_LOGIN_URL=... AUTH_STORAGE_STATE=... node scripts/auth-bootstrap-headful.mjs"), a copy-paste gitignore-guard bash snippet, a complete smoke-spec YAML block, and exact memory.write records. The gap keeping it from 5: Phase D says "write .claude/aw-targets/local.yml from the template" without ever showing or linking the local template (only the preview template gets an explicit path).

4 / 5

Workflow Clarity

Phases A–G are clearly sequenced with explicit validation (Phase C probe with expected outcomes, Phase E smoke verdict table with 'Loop back to Phase B with the specific failure'), an idempotency contract for re-runs, diff-before-overwrite confirmations, a stated preview variant path (A → B → C → D → G), and a Definition of done checklist — matching the anchor for explicit validation steps, feedback loops, and checklists.

5 / 5

Progressive Disclosure

Section structure and reference signaling are good, but no bundle files exist (no references/, scripts/, assets/, or templates/ directories) while the body links to ./templates/, ../templates/aw-tester.agent.md, and ../../../testing/ui-verify/... paths that don't resolve in the bundle, and ~105 lines of auth-template reference material are inlined in SKILL.md instead of living in the template files. This fits 'content that should be separate is inline' at 3 rather than the well-split structure of a 4.

3 / 5

Total

15

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: concrete, comprehensive actions, an explicit trigger clause, and a well-defined niche in the aw-tester/ui-verify toolchain. The only soft spot is keyword coverage, which is good but omits a few natural synonyms a user might say.

DimensionReasoningScore

Specificity

The description lists multiple concrete actions with comprehensive coverage: "detects auth strategy, captures storage state, writes .claude/aw-targets/<target>.yml, and validates with a smoke spec", plus "records the repo's UI surface" and "only re-prompt for what broke or changed" for re-runs. It covers first-run, re-run, and both targets (local/preview) — no gaps, matching the anchor for comprehensive coverage rather than the 'minor gaps' of a 4.

5 / 5

Completeness

It explicitly answers both questions: 'what' via the enumerated actions (scaffold, detect, capture, write, validate, record) and 'when' via the concrete trigger clause "Triggers on '/aw-setup', 'setup aw-tester', 'scaffold aw-target'" plus when-to-rerun guidance ("Re-runs detect the existing aw-target"). This matches the anchor 'clearly and explicitly answers both what AND when with concrete trigger phrases'; a 4 would require the 'when' to be less explicit.

5 / 5

Trigger Term Quality

Explicit triggers "Triggers on '/aw-setup', 'setup aw-tester', 'scaffold aw-target'" are natural phrases a user would say, plus domain keywords (auth strategy, storage state, preview, ui-verify). A few natural variations are missing (e.g. 'auth setup', 'configure test auth', 'login capture'), so it fits 'good keyword coverage; a few natural terms missing' rather than the comprehensive synonym coverage of a 5.

4 / 5

Distinctiveness Conflict Risk

It occupies a clear niche tied to a specific toolchain — "scaffolds a project's aw-tester aw-target", "ui-verify runs against", "the is-ui-diff gate" — with distinct trigger phrases unlikely to collide with any other skill. This is the 'clear niche with distinct triggers; minimal conflict risk' anchor, not merely 'mostly distinct' as at 4.

5 / 5

Total

19

/

20

Passed

Validation

68%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 11 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (638 lines); consider splitting into references/ and linking

Warning

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

relative_links

Relative link issues: 1 missing, 7 suspicious

Warning

referenced_paths_exist

Referenced path issues: 17 missing

Warning

Total

11

/

16

Passed

Repository
mthines/agent-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.