CtrlK
BlogDocsLog inGet started
Tessl Logo

sandbox-exec-fallback

Fallback pattern for executing Python code when execute_code_sandbox fails by writing to file and running via shell

58

Quality

73%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./benchmarks/gdpval/skills/sandbox-exec-fallback/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

67%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body delivers a clear, mostly executable recovery workflow with good structure and a verification step. It loses points for generic debugging tips and a weakly explicit feedback loop.

Suggestions

Remove or condense the "Tips" entries that restate basic programming advice (print debugging, exception handling) to improve token efficiency.

Make the Step 3 feedback loop explicit: "if errors appear, fix the script and re-run `python3 script.py` until it succeeds" before accessing output files.

Drop or shrink the "Why This Works" section, since the rationale (same Python environment, debuggable) is largely self-evident to Claude.

DimensionReasoningScore

Conciseness

Mostly efficient step-by-step guidance, but the "Why This Works" section and generic Tips ("Add explicit print statements for debugging", "Handle exceptions in your script") explain basic programming concepts Claude already knows and could be trimmed.

3 / 5

Actionability

Provides a concrete shell command ("python3 script.py"), a fully executable Python example, and specific tool parameters (path, content, timeout); the tool-invocation blocks are descriptive placeholders rather than copy-paste calls, leaving a minor gap.

4 / 5

Workflow Clarity

A clear four-step sequence (write → execute → verify → access) with a verification checkpoint in Step 3 and a mention of diagnosing/fixing, though the fix→re-run feedback loop is only implied rather than explicit.

4 / 5

Progressive Disclosure

Single self-contained file with well-organized sections and no external references; at ~80 lines it slightly exceeds the simple-skill threshold, keeping it just below the top band.

4 / 5

Total

15

/

20

Passed

Description

66%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description clearly states what the skill does and the specific failure condition that should trigger it, with low conflict risk. It is somewhat thin on action coverage and trigger variations, keeping it out of the top band.

Suggestions

Add 1-2 more concrete actions or outcomes (e.g., 'recover spreadsheet generation, data processing') to lift specificity.

Include a synonym or natural variant of the trigger (e.g., 'when sandbox execution errors out') to broaden trigger-term coverage.

DimensionReasoningScore

Specificity

Names the domain (Python code execution fallback) and two concrete actions ("writing to file and running via shell"), but coverage is not comprehensive — only the single recovery mechanism is described.

3 / 5

Completeness

Both "what" (fallback pattern for executing Python code via file + shell) and "when" (when execute_code_sandbox fails) are present and specific, though compressed into one sentence rather than a fuller multi-trigger "Use when..." clause.

4 / 5

Trigger Term Quality

The single trigger phrase "when execute_code_sandbox fails" is natural for this scenario, but no synonyms, variations, or related terms are included.

3 / 5

Distinctiveness Conflict Risk

The trigger is tied to a named tool's failure ("execute_code_sandbox fails"), giving it a clear niche with minimal overlap risk against other skills.

5 / 5

Total

15

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
HKUDS/OpenSpace
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.