CtrlK
BlogDocsLog inGet started
Tessl Logo

sandbox-fallback-python-e5d7ae

Fallback to run_shell with Python heredoc when execute_code_sandbox fails

57

Quality

72%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./benchmarks/gdpval/skills/sandbox-fallback-python-e5d7ae/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

75%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The content is highly actionable with complete runnable examples and a clear, well-organized fallback workflow including a verification checkpoint. Its main weaknesses are mild redundancy across the example sections and a comparison table that restates commonly known facts.

Suggestions

Collapse the basic heredoc, file-based, and PDF examples to avoid repeating the same cat > /tmp/script.py pattern; keep one canonical template and one fully-worked example.

Trim or remove the comparison table rows that state facts Claude already knows (isolation/containerization) and retain only the actionable deltas (timeout, file-path conventions).

Frame error handling as an explicit feedback loop: after running, check stderr; if imports/runtime errors appear, install missing libs or fix the script and re-run.

DimensionReasoningScore

Conciseness

Most content is task-specific and useful, but the comparison table restates facts Claude already knows (e.g., containerized isolation) and the basic, file-based, and PDF examples repeat the same cat-to-/tmp pattern, so it could be tightened.

3 / 5

Actionability

It provides multiple complete, copy-paste-ready executable examples covering short scripts, file-based scripts, library installation, and a full PDF generation plus verification flow across the common cases.

5 / 5

Workflow Clarity

Clear sequences appear in the file-based steps and the PDF example (write -> execute -> ls verify), plus a migration checklist, but no explicit validate -> fix -> retry feedback loop is framed as a loop.

4 / 5

Progressive Disclosure

The body is well-organized into clearly headed sections for a single-purpose skill with no need for external references, but at ~127 lines fully inline it has no overview-to-details split or signaled references.

4 / 5

Total

16

/

20

Passed

Description

55%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description clearly conveys a specific fallback action and an explicit trigger condition, giving it solid completeness and distinctiveness. It is held back by jargon-heavy trigger terms and a single narrow action rather than comprehensive coverage.

Suggestions

Add natural-language trigger phrases users would actually say (e.g., "Use when the code sandbox errors out or won't start, and you need to run Python with libraries like pandas or reportlab").

List a couple more concrete actions or covered scenarios (e.g., installing missing libraries, generating files/artifacts) to broaden specificity.

Keep the technical tool names but pair each with a plain-language synonym so the trigger is not purely jargon.

DimensionReasoningScore

Specificity

Names the domain and a concrete action ("Fallback to run_shell with Python heredoc") but lists only one fallback action rather than several, so it is not comprehensive.

3 / 5

Completeness

It states both what to do ("Fallback to run_shell with Python heredoc") and when ("when execute_code_sandbox fails"), but the single technical trigger condition falls short of comprehensive trigger phrases.

4 / 5

Trigger Term Quality

The keywords are mostly technical jargon ("execute_code_sandbox", "run_shell", "heredoc") with only "Python" and "fallback" as recognizable terms; it lacks the natural phrases a user would actually say.

2 / 5

Distinctiveness Conflict Risk

It carves a clear narrow niche (sandbox-failure fallback) with a distinctive trigger, with only minor overlap risk against general Python/shell execution skills.

4 / 5

Total

13

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
HKUDS/OpenSpace
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.