CtrlK
BlogDocsLog inGet started
Tessl Logo

skill-safety-reviewer

Review a skill-requested filesystem, command, network, secret, or destructive action against an explicit sandbox policy without executing it.

64

Quality

80%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./phases/13-tools-and-protocols/26-skill-permissions-sandboxes-and-trust/outputs/skill-safety-reviewer/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

93%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is lean, fully actionable, and well-structured with verified one-level-deep references, scoring at or near the top on every dimension. The only minor gap is the absence of an explicit error-recovery feedback loop around the review script run.

DimensionReasoningScore

Conciseness

The body is a lean ~11 lines with no padding and no explanation of concepts Claude already knows; it assumes competence and every line earns its place, matching the 'lean and efficient' anchor.

5 / 5

Actionability

Step 4 gives a fully copy-paste-ready, specific command ('python3 scripts/review_action.py --policy assets/sandbox-policy.json --request assets/example-request.json') and the other steps point to concrete real files, matching the 'fully executable; copy-paste ready' anchor.

5 / 5

Workflow Clarity

The five steps are clearly sequenced with real paths and step 5 returns a verdict (the validation output), but there is no explicit error-recovery feedback loop (e.g., what to do if the script errors or the verdict is ambiguous), so it sits at 'clear sequence with most checkpoints; minor validation gaps' rather than 5.

4 / 5

Progressive Disclosure

The SKILL.md is a concise overview that signals one-level-deep references to real bundle files (references/threat-model.md, scripts/review_action.py, assets/sandbox-policy.json, assets/example-request.json — all verified present), with content appropriately split and easy to navigate, matching the top anchor.

5 / 5

Total

19

/

20

Passed

Description

58%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific and clearly niche, but it lacks an explicit 'Use when ...' trigger clause, which caps completeness and leaves trigger-term quality mid-range. Adding a concrete usage trigger would lift the two weakest dimensions.

Suggestions

Add a 'Use when ...' clause naming natural trigger phrases (e.g., 'Use before a skill-driven workflow performs a filesystem write, shell command, network call, or destructive action').

Include common synonyms/extensions users might say (e.g., 'safe to run', 'is this command safe', 'sandbox check') to improve trigger-term coverage.

Keep the existing concrete action list but consider pairing it with the explicit when-guidance to reach a complete what+when statement.

DimensionReasoningScore

Specificity

The description enumerates concrete action categories ('filesystem, command, network, secret, or destructive action') and a precise action ('Review ... against an explicit sandbox policy without executing it'); it lists several specific actions rather than one comprehensive set, so it sits above the 3 anchor but below the comprehensive 5 anchor.

4 / 5

Completeness

It clearly states what the skill does ('Review a skill-requested ... action against an explicit sandbox policy without executing it') but provides no 'Use when ...' trigger clause, so per the rubric the missing explicit 'when' caps completeness at 3.

3 / 5

Trigger Term Quality

Terms like 'filesystem', 'command', 'network', 'secret', and 'destructive action' are relevant to the security domain, but 'sandbox policy' is jargon and common natural phrasings or synonyms a user would say are missing, matching the 'some relevant keywords but missing common variations' anchor.

3 / 5

Distinctiveness Conflict Risk

The niche is clear and specific (reviewing skill-requested actions against an explicit sandbox policy), giving it distinct triggers with only minor overlap risk against general security-review skills, fitting the 'mostly distinct; minor overlap risk' anchor rather than the fully distinct 5 anchor.

4 / 5

Total

14

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

Total

15

/

16

Passed

Repository
rohitg00/ai-engineering-from-scratch
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.