Build or adapt a bounded computer-use loop where Cua Driver observes and acts, TypeSafe Jev selects only from application-owned candidate IDs, and the caller validates and verifies every action. Use for the jev-use recipe or similar Jev integrations; do not use it to add model logic or credentials to Cua Driver.
Keep the decision layer above Cua Driver. Driver supplies observations and executes actions; the application constructs complete candidates; TypeSafe Jev returns one candidate ID. Never let Jev invent tool names, coordinates, refs, targets, delivery modes, or other arguments.
Use the example at libs/cua-driver/examples/jev-use/ as the runnable reference.
Keep TypeSafe request construction in the external Jev adapter rather than in
Driver or a Driver extension. The Python and TypeScript adapters must expose
equivalent mock and live behavior.
For a process boundary, use cua.jev_choice_request_v1 on stdin and require
cua.jev_choice_v1 on stdout. The request contains only a goal, capture ID,
compact regions, bounded history, and candidate IDs with descriptions; the
response contains only the selected ID, model identity, confidence, and
probabilities. Invoke the Python interpreter and absolute chooser path directly
without a shell.
For native desktop applications, use NativeAccessibilitySource and
cua.jev_choice_request_v2, which adds a per-candidate source (page, ax,
or visual), compact value-free elements, and optional progress counted
from the runner's own performed actions. Browser tasks keep sending v1.
Prefer browser DOM and semantic evidence. The optional visual adapter consumes
the public cua.visual_regions_v1 result only when Driver advertises both
parse_visual_regions and the capture-bound click.capture_id input.
Use the checked-in fixtures for deterministic development; do not add a model,
extension artifact, or Driver implementation detail to the recipe.
parse_visual_regions through the
current MCP tool inventory. Validate its versioned result, capture ID,
screenshot reference and dimensions, coordinate mapping, unique region IDs,
bounds, content, confidence, and ambiguity. Build a pixel action only with
the exact capture ID in the same click call. Otherwise reobserve or abstain.reobserve and abstain when
evidence can be stale, incomplete, or ambiguous.capture_id
when the current observation has validated visual evidence; do not send
extension internals or screenshot bytes.reobserve and abstain without inventing a mutation.capture_id or retry an expired, stale, or mismatched capture as
an unbound coordinate action.get_window_state call that returns the
tree and the screenshot together, so element tokens and capture_id
describe the same moment.native_roles.py / native_roles.ts, keyed by Driver's normalized_role.
Do not normalize roles in Driver.in_web_content elements, window chrome, and labels equal to the value.element_index. Cap at 24 action candidates plus reobserve and
abstain, and log how many were dropped.The deterministic mock path must work without TYPESAFE_API_KEY. For live Jev,
read the key from the process environment or a secure interactive prompt; never
put it in source, command arguments, logs, artifacts, or messages. Verify task
completion from an independent application postcondition rather than a model
answer, action response, or screenshot alone.
98e148b
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.