CtrlK
BlogDocsLog inGet started
Tessl Logo

sandbox-sdk

Build sandboxed applications for secure code execution. Load when building AI code execution, code interpreters, CI/CD systems, interactive dev environments, or executing untrusted code. Covers Sandbox SDK lifecycle, commands, files, code interpreter, and preview URLs. Biases towards retrieval from Cloudflare docs over pre-trained knowledge.

89

2.63x
Quality

85%

Does it follow best practices?

Impact

100%

2.63x

Average score across 3 eval scenarios

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

86%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The content is a well-structured, highly actionable overview: executable code for every common case, an explicit install-verification checkpoint, and clean progressive disclosure to two real one-level-deep reference files. The only notable weaknesses are a pinned version number outside a deprecation section and a missing post-setup validation step.

DimensionReasoningScore

Conciseness

The body is efficient and assumes Claude's competence — dense tables, tight code blocks, no concept over-explanation — matching the 4 anchor. It falls short of the 5 anchor because of minor time-sensitive detail (the pinned version "docker.io/cloudflare/sandbox:0.7.0") placed outside any 'old patterns'/deprecated section, which the guidelines penalize.

4 / 5

Actionability

Guidance is fully executable and copy-paste ready: install commands with a required check ("npm install @cloudflare/sandbox", "docker info"), the exact wrangler.jsonc block, the required "export { Sandbox }" entry snippet, and concrete API examples (exec, runCode with context reuse, file operations, exposePort) covering the common cases. This matches the 5 anchor.

5 / 5

Workflow Clarity

The sequence is clearly ordered (verify installation first, configure wrangler, export Sandbox, use the API, destroy for cleanup) and begins with an explicit validation checkpoint ("FIRST: Verify Installation" with docker info), plus cleanup guidance in the lifecycle and anti-patterns sections. It misses the 5 anchor because there is no explicit post-deploy/runtime verification step or error-recovery loop for the multi-step setup.

4 / 5

Progressive Disclosure

The body is a clear overview with common cases inline (Quick Reference table, core patterns) and the bulk detail correctly split into two well-signaled, one-level-deep references (references/api-quick-ref.md, references/examples.md), both of which exist and contain substantive content. This matches the 5 anchor; there is no nesting or buried reference.

5 / 5

Total

18

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description that clearly states what the skill does and gives an explicit, multi-trigger 'Load when' clause with natural phrases. Its main weaknesses are slightly nominal (rather than verb-based) capability listing and a couple of broad trigger terms that create minor conflict risk with adjacent skills.

DimensionReasoningScore

Specificity

The description lists several specific capability areas ("Sandbox SDK lifecycle, commands, files, code interpreter, and preview URLs") grounded in a concrete domain ("Build sandboxed applications for secure code execution"). It falls short of the 5 anchor because these are topical nouns rather than the multiple concrete verb-actions (e.g., "extract text, fill forms, merge documents") of a comprehensive listing, and ahead of 3 because coverage goes well beyond 1-2 actions.

4 / 5

Completeness

It explicitly answers both questions with concrete trigger phrases: the "what" via "Build sandboxed applications for secure code execution" plus the coverage list, and the "when" via the explicit "Load when building AI code execution, code interpreters, CI/CD systems, interactive dev environments, or executing untrusted code" clause. This matches the 5 anchor exactly; the 4 anchor applies only when the 'when' is less explicit.

5 / 5

Trigger Term Quality

Trigger terms are natural phrases users would say: "AI code execution, code interpreters, CI/CD systems, interactive dev environments, or executing untrusted code". A few natural synonyms/variations are missing (e.g., "run arbitrary code", "sandbox environment", file extensions), which keeps it below the comprehensive 5 anchor but clearly above the 3 anchor's partial coverage.

4 / 5

Distinctiveness Conflict Risk

The sandboxed-code-execution niche is mostly distinct, but "CI/CD systems" and "interactive dev environments" are broad enough that users building general CI/CD or dev-tooling projects could trigger this skill incorrectly, and the identifying vendor ("Cloudflare") appears only indirectly at the end. Minor overlap risk with closely related skills matches the 4 anchor rather than the minimal-conflict 5 anchor.

4 / 5

Total

17

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
openai/plugins
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.