CtrlK
BlogDocsLog inGet started
Tessl Logo

agent-sandbox

Agent skill for sandbox - invoke with $agent-sandbox

64

4.65x
Quality

47%

Does it follow best practices?

Impact

93%

4.65x

Average score across 3 eval scenarios

SecuritybySnyk

High

Do not use without reviewing

Fix and improve this skill with Tessl

tessl review fix ./.agents/skills/agent-sandbox/SKILL.md
SKILL.md
Quality
Evals
Security

Evaluation results

90%

68%

Automated Data Analysis Pipeline

Python sandbox lifecycle management

Criteria
Baseline
With context

Python template

0%

100%

Timeout set

0%

100%

Correct create function

0%

100%

File upload used

0%

100%

Correct execute function

0%

100%

capture_output enabled

0%

100%

Env vars for secrets

50%

100%

Sandbox cleanup

0%

100%

Status check

0%

0%

Error handling

100%

77%

Deployment order

100%

100%

91%

72%

Cross-Platform Application Smoke Test

Multi-environment sandbox orchestration and cleanup

Criteria
Baseline
With context

React template used

0%

100%

Python template used

0%

100%

Multiple sandbox_create calls

0%

100%

Correct create function

0%

100%

Status checked per environment

0%

0%

Execute in each environment

0%

100%

capture_output on executions

0%

100%

All sandboxes cleaned up

0%

100%

Cleanup despite failures

70%

100%

Timeout on each sandbox

0%

100%

Execution logging

100%

100%

Parallel or sequential orchestration

100%

100%

100%

81%

Node.js Microservice Test Harness

Node.js sandbox with package installation and status monitoring

Criteria
Baseline
With context

Node template

0%

100%

install_packages used

0%

100%

Correct create function

0%

100%

Correct execute function

0%

100%

capture_output enabled

0%

100%

Status monitoring

0%

100%

Timeout specified

0%

100%

Sandbox cleanup always runs

0%

100%

Error handling present

100%

100%

Execution logging

100%

100%

Language parameter

0%

100%

Repository
ruvnet/claude-flow
Evaluated
Agent
Claude Code
Model
Claude Sonnet 4.6

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.