CtrlK
BlogDocsLog inGet started
Tessl Logo

run-shell-python-file-io

Use run_shell with inline Python for reliable file I/O when execute_code_sandbox cannot access working directory files

59

Quality

74%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./benchmarks/gdpval/skills/run-shell-python-file-io/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

68%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is highly actionable with multiple complete, executable examples and a clear step sequence. Its weaknesses are mild verbosity from re-teaching basic Python file I/O and a missing validation checkpoint in the batch example.

Suggestions

Trim basic Python file I/O explanation and code comments that restate what the code already shows; keep only what illustrates the run_shell-vs-sandbox pattern.

Add an explicit verify/validation checkpoint (e.g., check expected output exists and is non-empty) to the batch processing example before declaring success.

DimensionReasoningScore

Conciseness

The skill is mostly efficient, but several code examples demonstrate basic Python file I/O patterns (open/read/write, os.getcwd) that Claude already knows, and prose like 'This provides reliable file I/O capabilities within the task workspace' is mild padding.

3 / 5

Actionability

Provides fully executable, copy-paste-ready bash/python examples covering the common cases: inline scripts via python3 -c, here-docs for complex logic, and batch glob processing, plus a troubleshooting table.

5 / 5

Workflow Clarity

Steps 1-4 are clearly sequenced and Step 4 provides verification, but the batch 'Processing Multiple Files' example overwrites files without any validation checkpoint; per the rubric cap, a batch/destructive workflow missing validation caps at 3.

3 / 5

Progressive Disclosure

A self-contained single-file skill with well-organized, clearly headed sections (When to Use, Pattern, How to Apply, Best Practices, Use Cases, Troubleshooting) and no nested references; good structure with minor organization gaps rather than a maximal split.

4 / 5

Total

15

/

20

Passed

Description

66%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is well-formed: it states a concrete mechanism and an explicit trigger condition tied to a specific tool limitation. It is distinct and unlikely to misfire, though its actions are generic and its trigger terms lean technical rather than natural.

Suggestions

Enumerate the concrete file operations supported (e.g., read, write, transform files) to raise specificity beyond generic 'file I/O'.

Add natural user-facing trigger phrases (e.g., 'Use when sandbox cannot read or write workspace files') alongside the technical condition.

DimensionReasoningScore

Specificity

Names the mechanism ('run_shell with inline Python') and a concrete but generic action ('reliable file I/O') without enumerating read/write/transform, so it lists domain plus 1-2 concrete actions rather than comprehensive coverage.

3 / 5

Completeness

Clearly states both what ('Use run_shell with inline Python for reliable file I/O') and when ('when execute_code_sandbox cannot access working directory files'); the 'when' is a concrete technical condition rather than a natural user-intent trigger phrase, keeping it just below a 5.

4 / 5

Trigger Term Quality

Includes relevant but technical/jargon keywords ('execute_code_sandbox', 'run_shell', 'file I/O', 'working directory files') with some natural phrasing, but misses common synonyms a user would naturally say.

3 / 5

Distinctiveness Conflict Risk

Targets a specific failure mode (sandbox cannot access working directory files) with named tools, carving a clear niche with distinct triggers and minimal conflict risk against general Python skills.

5 / 5

Total

15

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
HKUDS/OpenSpace
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.