CtrlK
BlogDocsLog inGet started
Tessl Logo

sandbox-bridge

Use when you need to exercise a real, running Sandbox deployment via HTTP — for example to validate SDK changes against a live container, reproduce a user-reported issue, or experiment with the API (including FUSE bucket mounts) without spinning up `wrangler dev`. Documents the Sandbox bridge worker reachable via `SANDBOX_WORKER_URL` + `SANDBOX_API_KEY` when the host injects them.

70

Quality

86%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

93%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An exemplary skill body: dense, executable, and free of padding, with copy-paste curl recipes for every core operation and disciplined deferral of full schemas to the bridge's own OpenAPI endpoint. The only meaningful gap is a missing validation checkpoint after sandbox creation, where a silent null ID would fail confusingly downstream.

DimensionReasoningScore

Conciseness

Lean and table-driven throughout: the SSE event table, the endpoint table, and compact curl snippets (including a dense one-line awk decode helper) assume Claude's competence and never re-teach known concepts (no explanation of what SSE or bearer auth is beyond what's needed). Every section carries bridge-specific facts Claude could not know otherwise, matching the level-5 anchor.

5 / 5

Actionability

Fully executable, copy-paste-ready curl commands cover the common cases end to end: create a sandbox, stream exec output with a decode pipeline, read/write files via PUT/GET, create/use/delete sessions, and destroy. Concrete details like the `argv` wrapping convention ("wrap shell snippets in [\"sh\",\"-lc\", ...]"), base64 decoding, and the Session-Id header leave nothing pseudocode-level.

5 / 5

Workflow Clarity

The Typical Flow is explicitly sequenced (create → exec → read/write files → destroy, with "Always clean up" and a -w "%{http_code}" check on DELETE), and error feedback is well covered (401 handling, the error/exit SSE events, and an Error Codes section). It falls short of 5 because creation is not validated: `SID=$(curl -s … | jq -r .id)` would silently set SID to null on failure (no `-sf` or exit-code check), so a failed create cascades into confusing downstream calls with no recovery guidance.

4 / 5

Progressive Disclosure

The body is a well-sectioned overview that defers depth exactly one level and signals it clearly: "The full, authoritative spec is served by the bridge itself" (/v1/openapi.json) and "If the API behaves unexpectedly, read those before guessing" pointing to the two source directories. Everything inlined (auth rules, event table, endpoint table, error codes) is compact enough to belong in the overview, and no nested references exist.

5 / 5

Total

19

/

20

Passed

Description

78%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description with explicit what/when structure, concrete use-case triggers, and clear differentiation from the local `wrangler dev` path. Its main deductions are the second-person phrasing (which the rubric penalizes on specificity) and reliance on jargon-heavy environment-variable names where more natural synonyms would improve trigger coverage.

Suggestions

Rewrite in third person to avoid the second-person specificity penalty — e.g. "Use when the user needs to exercise a real, running Sandbox deployment…" or "Exercises a real, running Sandbox deployment via HTTP. Use when…".

Broaden trigger coverage with natural synonyms users would actually say, such as "test against a live sandbox", "hit the sandbox API", or "mount an S3 bucket in a sandbox", instead of leaning on env-var names like SANDBOX_WORKER_URL.

Tighten the slightly generic "experiment with the API" trigger to something sandbox-specific (e.g. "try out sandbox API endpoints") to further reduce overlap with unrelated API-testing requests.

DimensionReasoningScore

Specificity

The description lists several concrete actions ("exercise a real, running Sandbox deployment via HTTP", "validate SDK changes against a live container", "reproduce a user-reported issue", "experiment with the API (including FUSE bucket mounts)"), which sits between the level-4 and level-5 anchors; however, the rubric's second-person penalty applies because it opens with "Use when you need to", reducing the base of 4 by 1. Not a 2 because the domain and actions are specific and multi-part rather than minimal or generic.

3 / 5

Completeness

Explicitly answers both questions: "Use when you need to exercise a real, running Sandbox deployment via HTTP — for example to validate SDK changes…, reproduce a user-reported issue, or experiment with the API" (when, with concrete triggers) and "Documents the Sandbox bridge worker reachable via SANDBOX_WORKER_URL + SANDBOX_API_KEY when the host injects them" (what). Matches the level-5 anchor with concrete trigger phrases; nothing is vague or merely implied.

5 / 5

Trigger Term Quality

Good natural keyword coverage — "validate SDK changes", "reproduce a user-reported issue", "FUSE bucket mounts", "wrangler dev", "Sandbox deployment" — phrases a user would plausibly say. Not 5 because common plain-language synonyms are missing and it leans on env-var jargon (SANDBOX_WORKER_URL, SANDBOX_API_KEY) that no user would naturally utter; not 3 because the coverage goes well beyond a single generic keyword.

4 / 5

Distinctiveness Conflict Risk

Clear niche — a hosted bridge worker driven over HTTP with bearer auth — and it explicitly disambiguates from the nearest competing skill ("without spinning up `wrangler dev`"). Minor overlap risk remains: "experiment with the API" is generic enough to fire for unrelated API-testing requests, so it is mostly rather than fully distinct.

4 / 5

Total

16

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
cloudflare/sandbox-sdk
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.