CtrlK
BlogDocsLog inGet started
Tessl Logo

cua-sandboxes

Create, use and clean up cua sandboxes (disposable Linux or macOS computers) locally or in the Cua cloud with the `cua` CLI or the cua SDK, and browse the web inside one. Use when a task needs an isolated machine to run code, test an app, drive a desktop GUI, browse or test a website, fill a web form, take a web page screenshot, or reproduce something without touching the user's own computer.

78

Quality

98%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

96%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A tightly written, highly actionable skill: complete commands and code across CLI, SDK, and MCP surfaces, a numbered browser workflow with an approval gate and re-read loops, and clear lifecycle and cleanup rules with zero padding. The only structural improvement available is splitting some inlined detail (SDK bindings, MCP browser workflow) into reference files for progressive disclosure.

Suggestions

Move the 'From code (Python)' SDK example and the TypeScript/Swift/Kotlin note into a references/ file (e.g. references/sdk.md), keeping a one-line pointer plus the single most common language inline in SKILL.md.

Split the detailed MCP browser-driving workflow (call_tool patterns, browser_* arguments) into references/browser.md and keep a short numbered summary with the key entry points in SKILL.md.

DimensionReasoningScore

Conciseness

The body is lean: command blocks with terse inline comments, one line of necessary context ('A sandbox is a disposable computer: a container or VM...'), and dense parameter documentation ('`--on` is where (`local`, `cloud`, `direct:<addr>`)...'). No concept Claude already knows is explained, and every token carries skill-specific information — matching 'Lean and efficient; assumes Claude's competence; every token earns its place'.

5 / 5

Actionability

Nearly everything is copy-paste ready: complete CLI commands for create/exec/shell/screenshot/vnc/cleanup, a fully executable Python example (including asyncio.run), and concrete MCP tool calls with argument names (e.g. `sandbox_create {"browser": true, "url": ...}`). Specific examples cover the common cases across CLI, SDK, and MCP, matching the top anchor.

5 / 5

Workflow Clarity

The lifecycle is clearly sequenced (Before you start → Create → Use → Clean up) with pre-flight checks (`auth status`, `runtime doctor`), documented error feedback ('An impossible combination fails with `invalid placement` and lists the valid values'), a re-read loop ('read again after every page change: refs are per snapshot'), and an explicit consent checkpoint before `teleport_browser_session`. Destructive operations here target disposable sandboxes — the isolation mechanism itself — so the missing-validation cap for destructive/batch operations does not apply; the content matches 'Clear sequence with explicit validation steps; feedback loops for error recovery'.

5 / 5

Progressive Disclosure

Sections are clear and flat (Create, Use, Clean up, From code, MCP, Browse the web, Rules) with no nested references and one clearly signaled cross-skill pointer ('see the gui-automation skill'), so it is well above anchor 3. It falls short of anchor 5 because there are no bundle reference files at all — everything, including the sizable SDK/MCP/browser detail, is inlined in a ~134-line body that could plausibly be split into one-level-deep references, and the under-50-line simple-skill exception does not apply.

4 / 5

Total

19

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An exemplary description: it states the full lifecycle capability (create, use, clean up, browse) with both interfaces and environments, and pairs it with an explicit, trigger-rich 'Use when' clause using natural user language. Third-person imperative voice is used correctly and there is no padding or over-claiming.

DimensionReasoningScore

Specificity

The description lists multiple concrete actions — 'Create, use and clean up cua sandboxes', 'browse the web inside one' — and covers both interfaces ('the `cua` CLI or the cua SDK') and both environments ('locally or in the Cua cloud'), giving comprehensive lifecycle coverage. It fits the anchor 'Lists multiple specific concrete actions; comprehensive coverage' rather than anchor 4, since no meaningful capability of the skill is left unmentioned.

5 / 5

Completeness

It explicitly answers both questions: the 'what' ('Create, use and clean up cua sandboxes... and browse the web inside one') and the 'when' via a concrete 'Use when a task needs an isolated machine...' clause with specific trigger scenarios. This is the anchor-5 pattern of clearly and explicitly answering both what AND when with concrete trigger phrases.

5 / 5

Trigger Term Quality

'run code, test an app, drive a desktop GUI, browse or test a website, fill a web form, take a web page screenshot, or reproduce something without touching the user's own computer' comprehensively covers natural phrases a user would actually say, including synonyms ('browse or test a website'). This matches the top anchor for natural-term coverage including variations.

5 / 5

Distinctiveness Conflict Risk

The sandbox/isolated-machine niche is distinct, and potentially overlapping terms (browsing, GUI, screenshots) are explicitly scoped to running 'inside one' sandbox, so conflict with generic browser or GUI-automation skills is minimal. It matches 'Clear niche with distinct triggers; minimal conflict risk'.

5 / 5

Total

20

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
trycua/cua
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.