CtrlK
BlogDocsLog inGet started
Tessl Logo

sandbox-failure-recovery

Recover from execute_code_sandbox failures by writing code to file and executing via run_shell

56

Quality

70%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./benchmarks/gdpval/skills/sandbox-failure-recovery/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

67%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The content is well-structured and highly actionable with a clear, verified workflow, but it carries some redundant examples and rationale that reduce conciseness, and lacks an explicit error-recovery feedback loop in the main workflow.

Suggestions

Remove the duplicated pandas/openpyxl example, keeping only the Complete Example or the Step 2 snippet, to tighten conciseness.

Promote error handling into an explicit workflow step (run -> if script errors, read traceback, fix, re-run) rather than relegating it to Best Practices.

Trim the "Why This Works" section to one line or fold it into the intro, since much of it states skill-specific rationale Claude can infer.

DimensionReasoningScore

Conciseness

The body is mostly efficient with concrete examples, but the "Why This Works" rationale and the duplicated pandas/openpyxl Excel example (Step 2 and the Complete Example) add padding that could be tightened, placing it at anchor 3.

3 / 5

Actionability

It provides concrete write_file/run_shell/list_dir/read_file invocations with parameters and a full copy-paste-ready Python example; not 5 because the tool-call blocks use a stylized pseudo-format and coverage centers on Excel/spreadsheet cases.

4 / 5

Workflow Clarity

Steps 1-4 are clearly sequenced with a verification checkpoint in Step 4; not 5 because there is no explicit failure-retry feedback loop embedded in the workflow itself (error handling lives only in Best Practices).

4 / 5

Progressive Disclosure

No bundle files exist and the ~127-line body is well-organized into clear sections with no inlined bulk material needing extraction; not 5 because it exceeds the under-50-line simple-skill threshold and contains duplicated content that could be trimmed.

4 / 5

Total

15

/

20

Passed

Description

57%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is concrete and clearly scoped to a distinctive niche, but it lacks an explicit "Use when..." trigger clause and richer keyword synonyms, which cap completeness and trigger-term quality at the midpoint.

Suggestions

Add an explicit 'Use when execute_code_sandbox fails, times out, or the sandbox is unstable' clause to lift completeness above 3.

Include natural user-facing synonyms such as 'sandbox error', 'code execution failed', or 'sandbox unavailable' to broaden trigger-term coverage.

List an additional concrete action or two (e.g. 'inspect error output', 're-run with longer timeout') to move specificity toward comprehensive coverage.

DimensionReasoningScore

Specificity

Names the domain (sandbox failure recovery) and two concrete actions ("writing code to file and executing via run_shell"), matching the anchor for domain plus 1-2 concrete actions; not 4 because it does not list several specific actions or variations.

3 / 5

Completeness

The "what" is clear (recover by writing to file and running via shell) and a "when" is present (sandbox failure), but there is no explicit standalone "Use when..." clause; per the rubric a missing explicit trigger clause caps completeness at 3.

3 / 5

Trigger Term Quality

Relevant keywords like "execute_code_sandbox failures" and "run_shell" appear, but there are no natural synonyms or variations a user might say (e.g. "sandbox crashed", "code execution failed"), so it sits at anchor 3.

3 / 5

Distinctiveness Conflict Risk

It targets a specific, distinctive trigger (execute_code_sandbox failure) with minimal overlap risk against unrelated skills, matching the "clear niche with distinct triggers" anchor.

5 / 5

Total

14

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
HKUDS/OpenSpace
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.