Use run_shell with inline Python as a reliable fallback when execute_code_sandbox or read_file fail
60
70%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
View guide
Critical
Do not install without reviewing
Fix and improve this skill with Tessl
tessl review fix ./benchmarks/gdpval/skills/run-shell-fallback/SKILL.mdSecurity
3 findings: 1 critical severity, 1 high severity, 1 medium severity. Installing this skill is not recommended: please review these findings carefully if you do intend to do so.
Detected high-risk code patterns in the skill content — including its prompts, tool definitions, and resources — such as data exfiltration, backdoors, remote code execution, credential theft, system compromise, supply chain attacks, and obfuscation techniques.
This document explicitly instructs bypassing sandbox protections by using run_shell to execute arbitrary shell/Python code and read/write filesystem files, enabling potential data exfiltration and local compromise.
The skill handles credentials insecurely by requiring the agent to include secret values verbatim in its generated output. This exposes credentials in the agent’s context and conversation history, creating a risk of data exfiltration.
The skill's examples and patterns instruct the agent to use run_shell to read files and print their contents (e.g., cat, json.load(...) and print(data['key'])), which would cause any secrets present in those files to be output verbatim by the LLM.
The skill prompts the agent to compromise the security or integrity of the user’s machine by modifying system-level services or configurations, such as obtaining elevated privileges, altering startup scripts, or changing system-wide settings.
The skill explicitly instructs bypassing the sandbox to run shell commands (run_shell) that read arbitrary files and write output files, enabling modification of the host filesystem and access to sensitive paths even though it does not explicitly request sudo or privileged operations.
c5a9c4b
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.