CtrlK
BlogDocsLog inGet started
Tessl Logo

sandbox-manager

Inspect, upgrade, or restart the sandbox this Agent is running in. Use when the user asks about this Agent's sandbox status or Agent image version, or explicitly asks to upgrade or restart its current environment.

73

Quality

90%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is actionable and well-sequenced with strong validation for its destructive operations. Minor conciseness and structure gains remain, mostly from duplicated scheduling prose.

Suggestions

Collapse the repeated 'reply now / verify after recovery' guidance — stated once as a shared rule rather than re-explained after both the upgrade and restart scheduled-result blocks — to tighten conciseness.

Consider moving the long sample `result.data` JSON blocks into a reference file and keeping only one representative example inline, improving progressive disclosure for a skill that is over 50 lines.

DimensionReasoningScore

Conciseness

The body is efficient with executable snippets and no padding of concepts Claude already knows, but the scheduled-result blockquote ('Reply to the user now and do not call more tools...') is restated in the prose immediately after, a minor trim opportunity; fits score-4 better than 5.

4 / 5

Actionability

Each operation ships a complete, copy-paste `run_sdk_snippet` example with error handling and result inspection, covering the common cases fully, matching the score-5 anchor.

5 / 5

Workflow Clarity

Destructive operations carry explicit validation checkpoints — the `if not result.ok` pattern, an explicit-user-intent gate, post-recovery `get_sandbox_info` verification, and a no-auto-retry rule — matching the score-5 feedback-loop anchor.

5 / 5

Progressive Disclosure

No bundle files exist and the content is self-contained with well-organized section headers, so structure is good; a few inline example outputs could be condensed, leaving it just below the score-5 ideal.

4 / 5

Total

18

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is strong: it states concrete capabilities, gives explicit trigger guidance, and is clearly distinct from other skills. The only minor gap is coverage of a few natural phrasings users might use.

DimensionReasoningScore

Specificity

Names the domain (sandbox) and three concrete actions — 'Inspect, upgrade, or restart' — covering the full capability set comprehensively, matching the score-5 anchor.

5 / 5

Completeness

Explicitly answers both 'what' (inspect/upgrade/restart the current sandbox) and 'when' via a concrete 'Use when the user asks about...' clause, matching the score-5 anchor.

5 / 5

Trigger Term Quality

Natural terms like 'sandbox status', 'Agent image version', 'upgrade', and 'restart' are present and user-likely, but a few synonyms (e.g. 'update my agent', 'what version is the sandbox') are missing, fitting the score-4 anchor better than 5.

4 / 5

Distinctiveness Conflict Risk

The phrase 'the sandbox this Agent is running in' carves a clear niche with distinct triggers and minimal overlap with other skills, matching the score-5 anchor.

5 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
dtyq/magic
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.