CtrlK
BlogDocsLog inGet started
Tessl Logo

cua-driver

Drive a native GUI app (macOS, Windows, Linux) via the cua-driver CLI (default) or MCP server; snapshot its accessibility tree, act through snapshot-bound element tokens, native menu paths, exact window geometry, or pixel coordinates, and verify from fresh state. Use when the user asks you to operate, drive, automate, or perform a GUI task in a real application on the host, or to continue, resume, or recall recent Cua activity.

76

Quality

98%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

96%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A model of lean, high-density skill authoring: the body is pure routing (what to do, in what order, with which exact command or tool, and what to do when it fails) and delegates all depth to per-topic files via clearly signaled, anchor-linked references. The one material defect is bundle completeness: every file the body routes to is absent from the evaluated bundle, which leaves the otherwise excellent navigation unresolvable.

Suggestions

Ship (or restore) the referenced bundle files next to SKILL.md — WORKFLOW.md, RUNTIME.md, BROWSER.md, RECORDING.md, EMBEDDING.md, and the current-host platform guide — so the body's routing table and failure-map deep links resolve; without them every 'Read when needed' pointer is a dead link.

If this bundle is intentionally host-filtered, state that in the body: the current note only covers platform files ('Other platform files may be absent from a host-filtered installation'), but WORKFLOW.md and RUNTIME.md are core dependencies and their absence is unexplained.

Consider one short inline example of a minimal observe→act→verify exchange (e.g., a two-line get_window_state → click(element_token) sequence) so the mainline loop is executable even when WORKFLOW.md is unavailable in a filtered installation.

DimensionReasoningScore

Conciseness

The ~76-line body is entirely operational with zero padding: a one-line north star ('Operate one exact target, observe its state, act once, and verify the user's postcondition'), a goal→tool routing table, seven dense rules, and a symptom→next-step failure map. It explains nothing Claude already knows, matching the lean/every-token-earns-its-place anchor rather than the 4 anchor's 'minor instances of over-explanation'.

5 / 5

Actionability

For an instruction-only skill, guidance is concrete throughout: exact commands ('cua-driver --version, status, doctor, describe <tool>'), argument-shaped tool calls ('get_window_state({pid, window_id})', 'verify_state({pid, window_id, expect})', 'target:{kind:"desktop",display_id:"primary"}'), and explicit sequences ('start_recording → actions → stop_recording'). Specific commands cover the common cases across observe, act, verify, desktop, browser, and recording, satisfying the fully-executable anchor; it sits above the 4 anchor because no mainline step is left at the hint level.

5 / 5

Workflow Clarity

The sequence (observe → act once → verify → stop after proof) is stated in the summary, Rules #2, and the Act table's row order, with explicit validation checkpoints ('verify_state({pid, window_id, expect}) or a fresh snapshot', 'effect:"unverifiable" and a successful exit are not task success'). Feedback loops are present in the Failure map (e.g., 'Text did not visibly change → Reobserve before retrying') and the history section is a fully-branched decision flow, matching the anchor for clear sequence, explicit validation, and error-recovery loops.

5 / 5

Progressive Disclosure

The on-paper structure is anchor-5 quality: a lean overview delegating all detail one level deep, a 'Read when needed' column, deep links to specific anchors (e.g., 'RUNTIME.md#preflight-and-transport', 'WORKFLOW.md#act-once'), and an annotated References section with 'Load on demand; do not reabsorb these into this file'. However, none of the eight referenced files (WORKFLOW.md, RUNTIME.md, MACOS.md, WINDOWS.md, LINUX.md, BROWSER.md, RECORDING.md, EMBEDDING.md) exist in the provided bundle — there is no references/ directory and no such files beside SKILL.md — so every navigation target dangles in practice.

4 / 5

Total

19

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An exemplary description: concrete what (snapshot/act/verify with named actuation modalities and transports) paired with an explicit, synonym-rich when-clause covering both new GUI tasks and continuation of prior Cua activity. Voice is consistent with the third-person/imperative convention and the niche is sharply distinguished from adjacent skills.

DimensionReasoningScore

Specificity

The description lists multiple concrete actions covering the full operate loop — 'snapshot its accessibility tree, act through snapshot-bound element tokens, native menu paths, exact window geometry, or pixel coordinates, and verify from fresh state' — plus transport choice ('via the cua-driver CLI (default) or MCP server') and platform coverage. This matches the anchor for multiple specific concrete actions with comprehensive coverage; not the 4 anchor, which requires minor coverage gaps.

5 / 5

Completeness

It explicitly answers both questions: 'what' is the concrete snapshot/act/verify capability set, and 'when' is the explicit clause 'Use when the user asks you to operate, drive, automate, or perform a GUI task... or to continue, resume, or recall recent Cua activity' with concrete trigger phrases. This is the anchor-5 pattern verbatim in structure; the 4 anchor requires a 'when' that could be more explicit, which does not apply.

5 / 5

Trigger Term Quality

It includes natural synonym clusters users would actually say: 'operate, drive, automate, or perform a GUI task in a real application on the host' and 'continue, resume, or recall recent Cua activity'. Coverage spans the domain's common natural phrasings, matching the comprehensive-synonyms anchor rather than the 4 anchor ('a few natural terms missing').

5 / 5

Distinctiveness Conflict Risk

The niche is clearly bounded to driving native GUI apps via accessibility-tree snapshots ('native GUI app (macOS, Windows, Linux)', 'snapshot-bound element tokens', 'cua-driver'), which distinguishes it from browser-automation, clipboard, and shell skills, and the triggers name host GUI operation specifically. Minimal conflict risk matches the 5 anchor; it is not merely 'mostly distinct' as in the 4 anchor.

5 / 5

Total

20

/

20

Passed

Validation

75%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 12 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

relative_links

Relative link issues: 23 missing

Warning

Total

12

/

16

Passed

Repository
trycua/cua
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.