CtrlK
BlogDocsLog inGet started
Tessl Logo

run-shell-fallback

Use run_shell with inline Python as a reliable fallback when execute_code_sandbox or read_file fail

55

Quality

69%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Critical

Do not install without reviewing

Fix and improve this skill with Tessl

tessl review fix ./benchmarks/gdpval/skills/run-shell-fallback/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

61%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The content is highly actionable with executable code and a clear decision flow, organized into clean self-contained sections. It loses points for a redundant Examples section, missing validation checkpoints in the workflow, and an un-demonstrated error-handling best practice.

Suggestions

Collapse the Examples section or fold its distinct cases (CSV, config.json) into How to Implement to remove redundancy.

Add a validation checkpoint to the Decision Flow, e.g. "Check run_shell exit code / stderr before continuing".

Show one inline Python example with a try/except block to back up Best Practice #5.

DimensionReasoningScore

Conciseness

The body avoids explaining concepts Claude already knows, but the standalone Examples section largely re-demonstrates patterns already shown in How to Implement (JSON parse, multi-line processing, file write), so it could be tightened.

3 / 5

Actionability

Code blocks are concrete and copy-paste-ready across one-liners, heredoc, file reads, and inline writes; the minor gap is that Best Practice #5 recommends try/except yet no example demonstrates it.

4 / 5

Workflow Clarity

The Decision Flow gives a clear numbered sequence, but there are no validation checkpoints (e.g. verify run_shell succeeded, check exit code) and some variants write files, which is mildly risky without a verify step.

3 / 5

Progressive Disclosure

The body is well-organized into clearly headed sections (When to Use, How to Implement, Best Practices, Examples, Limitations) and needs no external references; redundancy between Examples and How to Implement keeps it just below the top band.

4 / 5

Total

14

/

20

Passed

Description

62%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description cleanly states what the skill does and when to use it with a specific failure trigger, giving it strong activation guidance. It is held below the top band by limited action enumeration and missing common trigger variations like "unknown error" and "timeout".

Suggestions

Add the concrete failure signatures users actually encounter, e.g. "when execute_code_sandbox or read_file fail with 'unknown error' or timeout".

Briefly enumerate the fallback variants (one-liner `-c`, heredoc for multi-line, file read vs code execution) so the description reflects the skill's full scope.

DimensionReasoningScore

Specificity

Names the tool and one concrete action ("Use run_shell with inline Python as a reliable fallback"), but coverage is not comprehensive — it does not enumerate the variants (one-liners, heredoc, file reads vs code execution) that the body actually covers.

3 / 5

Completeness

It explicitly answers both "what" (use run_shell with inline Python as a fallback) and "when" (when execute_code_sandbox or read_file fail); the "when" is concrete but a single trigger without the variations that would push it to 5.

4 / 5

Trigger Term Quality

Relevant keywords are present (run_shell, inline Python, execute_code_sandbox, read_file, fallback), but it omits common natural variations a user would actually say such as "unknown error" and "timeout", which the body treats as the real signatures.

3 / 5

Distinctiveness Conflict Risk

The failure-triggered niche is distinct and unlikely to fire for unrelated skills, with only minor overlap risk against general shell-execution or Python skills.

4 / 5

Total

14

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
HKUDS/OpenSpace
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.