CtrlK
BlogDocsLog inGet started
Tessl Logo

gui-automation

Use when you need to visually interact with a GUI: test buttons, fill forms, verify visual layouts, fuzz web pages, automate user flows, take screenshots, or perform end-to-end QA on any application. Works on cloud VMs, Docker containers, local machines, and sandboxes. Needs the `cua` CLI (curl -fsSL https://cua.ai/install.sh | sh).

74

Quality

91%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

92%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

Excellent content: fully executable commands throughout, an explicit Look→Act→Verify loop with staleness warnings as validation, and a clean split between the in-body overview and the one-level-deep command reference. The only improvement space is minor redundancy between the Workflow/Scenarios and Setup/Providers sections.

DimensionReasoningScore

Conciseness

The body is lean and command-first with no explanations of concepts Claude already knows — the zoom section's note about window-relative coordinates is genuinely non-obvious tool behavior. Minor redundancy keeps it below the top anchor: the "Click a button" scenario repeats the exact commands already shown in the Workflow section, and the Providers table re-lists the `cua do switch` invocations from Setup.

4 / 5

Actionability

Everything is copy-paste-ready executable bash with comments: install check, target switching, screenshot-click-verify loops, form fill with tab navigation, file upload, zoom/unzoom for precision, drag and drop, fuzz payloads, and trajectory commands. The quick-reference table covers all common verbs. The only placeholders (coordinates like `450 280`) are inherent to the skill, and the workflow explains they come from the preceding screenshot.

5 / 5

Workflow Clarity

The core loop is explicit and named — "Look → Act → Verify: repeat until done" — with a strong validation callout: "Re-screenshot after every UI change: coordinates go stale when the screen changes." Every scenario ends with a verification screenshot, the fuzz scenario tells what to check for ("errors, crashes, unexpected behavior"), and the zoom path includes unzoom-restore steps. This matches the anchor with explicit validation steps and feedback loops.

5 / 5

Progressive Disclosure

Scored against the actual bundle: the body is an overview plus short actionable scenarios, and the full argument syntax is correctly delegated to references/command-reference.md — a real file, referenced exactly once, one level deep, clearly signaled at the end of the body. The inline quick-reference table is a compact cheat sheet, not bulk detail that belongs in the reference. Navigation is easy with consistent section headers.

5 / 5

Total

19

/

20

Passed

Description

91%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: comprehensive concrete capabilities, explicit trigger phrasing, and useful environment/install details. Its only deductions are the second-person voice (penalized on specificity per rubric) and slight overlap risk with browser/web-testing skills.

DimensionReasoningScore

Specificity

The description lists many concrete actions — "test buttons, fill forms, verify visual layouts, fuzz web pages, automate user flows, take screenshots, or perform end-to-end QA" — plus deployment targets ("cloud VMs, Docker containers, local machines, and sandboxes") and the install command, which would merit a 5. However, it opens with second-person voice ("Use when you need to..."), and the rubric penalizes non-third-person phrasing by reducing specificity by 1, landing at 4 — several specific actions with minor gaps, not the comprehensive-anchor exemplar.

4 / 5

Completeness

It explicitly answers both: what — visually interact with a GUI (test, fill, verify, fuzz, screenshot, QA); when — an explicit "Use when you need to visually interact with a GUI" clause with concrete trigger phrases. It also adds environment and dependency details, matching the top anchor.

5 / 5

Trigger Term Quality

Trigger phrases are the natural words a user would say: "test buttons", "fill forms", "fuzz web pages", "automate user flows", "take screenshots", "end-to-end QA", "visually interact with a GUI". Coverage spans synonyms across testing, QA, and automation vocabulary; a user asking for any of these would naturally match this description.

5 / 5

Distinctiveness Conflict Risk

The core niche — visual/screenshot-driven GUI interaction on a real machine — is clear and distinct. Minor overlap risk remains with closely related browser-automation or web-testing skills, since "fill forms", "fuzz web pages", and "end-to-end QA" are triggers those skills could also claim, so it fits the "mostly distinct; minor overlap risk" anchor rather than the minimal-conflict anchor.

4 / 5

Total

18

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
trycua/cua
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.