CtrlK
BlogDocsLog inGet started
Tessl Logo

iii-sandbox

Ephemeral microVM sandboxes for running untrusted or agent-generated code in isolation — a one-call run path, a create/exec/stop lifecycle, and a set of filesystem operations.

62

Quality

72%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./crates/iii-worker/src/sandbox_daemon/skills/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

75%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The content is well-structured, lean, and actionable, with a clear two-path workflow (one-call run vs. create/exec/stop lifecycle) and a real error-recovery feedback loop. It consistently lands at anchor 4: strong with only minor gaps in conciseness, inline copy-paste examples, explicit validation checkpoints, and file-splitting.

DimensionReasoningScore

Conciseness

The body is mostly lean and assumes competence (no basic-concept padding), but the dense localhost→host rewrite paragraph and a few explanatory passages could be trimmed, fitting the "efficient; minor instances of over-explanation" anchor 4 rather than the fully lean anchor 5.

4 / 5

Actionability

Concrete guidance is present — the executable `engine::functions::info { function_id: "sandbox::<fn>" }`, a named function map, and specific gotchas (env at create-time, backgrounded exec) — but full copy-paste call examples are deliberately deferred to the engine contract, leaving minor gaps at anchor 4 rather than the copy-paste-ready anchor 5.

4 / 5

Workflow Clarity

The create → exec/fs::* → stop lifecycle is clearly sequenced and a feedback loop exists ("read it and apply the fix before retrying") plus recovery (replace the sandbox), but there are no explicit numbered validation checkpoints or checklists, matching anchor 4; the destructive-operation cap to 3 does not apply because validation guidance is present.

4 / 5

Progressive Disclosure

Sections are well-organized and the authoritative function contracts are deferred to the engine as a clearly signaled one-level reference, but the skill is a single ~67-line file with an inline function reference list and no content split across files, fitting anchor 4 rather than the fully-split anchor 5.

4 / 5

Total

16

/

20

Passed

Description

70%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific and clearly distinct, naming several concrete capabilities around an isolated-code-execution niche. Its main weakness is the missing explicit "Use when…" trigger guidance, which leaves the "when" only weakly implied and caps completeness.

Suggestions

Append an explicit "Use when…" clause, e.g. "Use when you need to run untrusted or LLM-generated code in isolation, do a one-shot code-in/stdout-out run, or manipulate files inside an isolated filesystem." — this lifts completeness above the 3 cap.

Enumerate the filesystem operations briefly (read, write, search, edit) instead of "a set of filesystem operations" to move specificity toward anchor 5.

Add a couple of natural synonyms or concrete trigger phrases (e.g. "sandbox", "execute code safely", "isolated code runner") to broaden natural-term coverage beyond the current jargon-leaning set.

DimensionReasoningScore

Specificity

Quotes "a one-call run path", "a create/exec/stop lifecycle", and "a set of filesystem operations" — several concrete capabilities are named, but "a set of filesystem operations" is left unenumerated, a minor coverage gap matching anchor 4 rather than the comprehensive anchor 5.

4 / 5

Completeness

The "what" is clearly stated ("Ephemeral microVM sandboxes for running untrusted or agent-generated code in isolation") but there is no explicit "Use when…" trigger clause, so "when" is only weakly implied; per the judging guideline this caps completeness at anchor 3.

3 / 5

Trigger Term Quality

Natural terms like "sandbox", "untrusted", "agent-generated code", and "isolation" appear alongside jargon ("microVM", "one-call run path"); coverage is good but lacks the synonym/extension breadth of anchor 5, fitting anchor 4.

4 / 5

Distinctiveness Conflict Risk

The niche — isolated microVM sandboxes for untrusted/agent-generated code — is clearly distinct with trigger terms unlikely to fire for unrelated skills, matching the "clear niche with minimal conflict risk" anchor 5.

5 / 5

Total

16

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
iii-hq/iii
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.