CtrlK
BlogDocsLog inGet started
Tessl Logo

agent-sandbox

Agent skill for sandbox - invoke with $agent-sandbox

61

4.65x
Quality

43%

Does it follow best practices?

Impact

93%

4.65x

Average score across 3 eval scenarios

SecuritybySnyk

High

Do not use without reviewing

Fix and improve this skill with Tessl

tessl review fix ./.agents/skills/agent-sandbox/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

57%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a serviceable agent prompt with a genuinely useful MCP toolkit reference, but it is undermined by a duplicated stray frontmatter block, a prompt-style voice rather than skill documentation, and no validation or cleanup checkpoints in its workflow. Structure is flat — no headers, no external references — leaving it a monolithic inline prompt.

Suggestions

Remove the second '---' frontmatter block (flow-nexus-sandbox) from the body — it is a parsing hazard and duplicates metadata that belongs in the single top frontmatter.

Add verification before destructive operations, e.g., check `sandbox_status` output and confirm with the user before calling `sandbox_delete`, and validate execution results after `sandbox_execute`.

Convert the body to skill-document format with markdown headers (## Toolkit, ## Templates, ## Workflow) and move the template reference and API details into `references/` files, keeping SKILL.md a lean overview.

DimensionReasoningScore

Conciseness

The toolkit block is dense and useful, but "Quality standards" ("Implement proper error handling and logging", "Secure environment variable management") and the closing paragraph are generic filler that teaches nothing. Mostly efficient with some padding — anchor 3, not 2, because the core toolkit and template sections do earn their tokens.

3 / 5

Actionability

The toolkit block shows each MCP tool call with full parameter signatures (template, env_vars, install_packages, timeout, etc.), which is concrete, actionable guidance. Minor gaps keep it at anchor 4 rather than 5: placeholders like "sandbox_id" and "$app$config.json" are unfilled/typo'd, and there is no example of consuming the create response.

4 / 5

Workflow Clarity

The 6-step "deployment approach" gives a clear sequence but steps are high-level ("Track resource usage and execution metrics") with no validation checkpoints or commands. Because sandbox deletion is a destructive operation with no verification step, the workflow-clarity cap of 3 applies — anchor 3 is the best fit.

3 / 5

Progressive Disclosure

Labeled sections (toolkit, templates, approach, standards) provide some structure, but there are no markdown headers, no reference files at all, and the ~75-line prompt inlines everything including template/API detail that belongs in separate files. Anchor 3 ("Some structure but could be better organized; content that should be separate is inline").

3 / 5

Total

13

/

20

Passed

Description

28%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is a placeholder-grade label: it identifies the domain but says nothing about what the skill does, when to invoke it, or which tools it manages. It would rarely win a trigger match against a better-described sandbox skill.

Suggestions

State concrete capabilities in third person, e.g., "Creates, configures, and manages E2B sandboxes; executes code in isolation; handles file uploads and sandbox lifecycle" — the second frontmatter block in the body already contains a usable version of exactly this text.

Add an explicit trigger clause: "Use when the user mentions sandboxes, E2B, isolated execution environments, or running untrusted code safely."

Include natural keyword variations (E2B, sandbox, execution environment, run code in isolation) so the skill matches how users actually phrase these requests.

DimensionReasoningScore

Specificity

The description names the domain ("sandbox") but lists no concrete actions whatsoever — "Agent skill for sandbox" is a label, not a capability statement. It matches anchor 2 ("Names the domain but actions are minimal or generic") and falls short of anchor 3, which requires 1-2 concrete actions.

2 / 5

Completeness

It has a vague 'what' ("Agent skill for sandbox") and no 'when' guidance at all — an exact match for anchor 2. The missing "Use when..." clause caps completeness at 3 anyway, and the vagueness of the 'what' places it below that.

2 / 5

Trigger Term Quality

Only "sandbox" and "invoke" appear as keywords, with no natural user phrases or variations (e.g., E2B, execution environment, isolated testing, run code). This fits anchor 2 ("One or two generic keywords; missing the natural phrases users say") rather than anchor 3, since there is no synonym or variation coverage at all.

2 / 5

Distinctiveness Conflict Risk

"sandbox" points at a recognizable niche and the "$agent-sandbox" identifier is somewhat distinct, but the phrasing is broad enough to overlap with other agent or sandbox-related skills. Fits anchor 3 ("Somewhat specific but could still overlap with similar skills") more than anchor 2, since the domain term is not purely generic.

3 / 5

Total

9

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
ruvnet/ruflo
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.