Create new skills, modify and improve existing skills, and measure skill performance. Use when users want to create a skill from scratch, edit, or optimize an existing skill, run evals to test a skill, benchmark skill performance with variance analysis, or optimize a skill's description for better triggering accuracy.
67
81%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
View guide
Low
Low-risk findings worth noting
Low
Low-risk findings.
2 low severity findings. Worth noting, but not necessarily harmful.
The skill exposes the agent to untrusted, user-generated content from public third-party sources, creating a risk of indirect prompt injection. This includes browsing arbitrary URLs, reading social media posts or forum comments, and analyzing content from unknown websites.
The required workflow uses `eval-viewer/generate_review.py` to read and embed arbitrary text/markdown from runtime-produced run files like `outputs/transcript.md` / `user_notes.md` / `grading.json` into the served HTML (and thus into the context shown to the reviewing model/user), where those files can contain outsider-authored free text from the executed eval prompts and user feedback.
The skill fetches instructions or code from an external URL at runtime, and the fetched content directly controls the agent’s prompts or executes code. This dynamic dependency allows the external source to modify the agent’s behavior without any changes to the skill itself.
The eval viewer HTML (eval-viewer/viewer.html) includes a runtime-loaded remote JavaScript library (https://cdn.sheetjs.com/xlsx-0.20.3/package/dist/xlsx.full.min.js) which will be fetched and executed in the user's browser when the viewer is opened, representing a runtime external dependency that executes remote code.
f618458
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.