Drive a native GUI app (macOS, Windows, Linux) via the cua-driver CLI (default) or MCP server; snapshot its accessibility tree, act through snapshot-bound element tokens, native menu paths, exact window geometry, or pixel coordinates, and verify from fresh state. Use when the user asks you to operate, drive, automate, or perform a GUI task in a real application on the host, or to continue, resume, or recall recent Cua activity.
72
92%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
View guide
Passed
No findings from the security scan
The canonical home for this skill is cua-driver in trycua/cua
Operate one exact target, observe its state, act once, and verify the user's postcondition.
| Goal | Tool or command | Read when needed |
|---|---|---|
| Check installation and capabilities | cua-driver --version, status, doctor, describe <tool>; MCP tools/list | Runtime |
| Find or open the requested app | list_apps, list_windows, launch_app | Current platform guide below |
| Observe one window | get_window_state({pid, window_id}) | Workflow |
| Act on a control | click / type_text with a fresh element_token and exact window target | Workflow |
| Use pixels when semantics cannot reach it | Fresh target screenshot, then x,y on the same target | Workflow |
| Verify the outcome | verify_state({pid, window_id, expect}) or a fresh snapshot read by the agent | Workflow |
| Operate the authorized desktop | get_desktop_state → input with target:{kind:"desktop",display_id:"primary"} → get_desktop_state | Workflow, Linux on Wayland |
| Drive supported browser page content | get_browser_state → typed browser action → fresh state | Browser |
| Record an explicitly requested run | start_recording → actions → stop_recording; verify artifacts | Recording |
| Finish | Stop after proof; end_session for this run, not cua-driver stop on a shared service | Runtime |
Use Cua when the outcome lives in an application's UI/window state or the user asks to operate that GUI. Honor a requested interaction method: GUI-only excludes application APIs, DOM/CDP, direct clipboard APIs, and shell mutations unless the user permits them.
Check the installed version and advertised schema before using unfamiliar parameters. This pack's version identifies its source release, not the running daemon. Do not upgrade software, change permission profiles, or reinstall skills merely to make a recipe work.
effect:"unverifiable" and a successful exit are not task success; never replay a partial, canceled, or unknown action blindly.invalidated_snapshot_ids; act with element_token.| Symptom | Next step |
|---|---|
| Missing binary, mismatched daemon, unknown tool/field | Runtime preflight |
| Stale token or ambiguous window | Refresh list_windows / get_window_state; choose the intended live target |
| Large or sparse tree | Bounded observation |
surface_identity_unproven or screenshot permission wait | Wayland capture recovery |
background_unavailable | Verify current state; ask before foreground/desktop control if not already authorized |
| Text did not visibly change | Reobserve before retrying; text and value semantics |
| Browser setup, binding, or ref refused | Browser recovery |
When both history_status and history_query are advertised and the user asks
to continue, resume, or recall prior Cua work, call history_status first. If
history is healthy and access is admitted, make one bounded initial
history_query before broad application or window discovery. Treat returned
metadata only as a lead and verify current state through the least intrusive
appropriate source. Content, geometry, arguments, results, and user intent
omitted from the metadata remain unknown.
Make another bounded query only when the initial slice exposes a relevant session or sequence boundary; never broaden a query to reconstruct excluded fields.
Continue without history when either tool is absent, access is denied, the query is empty, or history is unhealthy. Do not query history for unrelated tasks merely because the tools are advertised, and never mutate history lifecycle or settings.
Load on demand; do not reabsorb these into this file:
98e148b
Canonical home
since Jul 24, 2026
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.