CtrlK
BlogDocsLog inGet started
Tessl Logo

llm-security

Use for authorized security assessment of LLM applications and AI agents, including prompt injection, tool abuse, RAG exposure, memory poisoning, and model supply-chain risks.

76

Quality

93%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Critical

Do not install without reviewing

SKILL.md
Quality
Evals
Security

Quality

Content

87%Weight 40%Scale 1-3

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A dense, highly actionable security-testing skill with concrete probes, tooling, and a clean reference split. The main gap is the absence of explicit validation/feedback checkpoints within the batch-oriented testing workflow.

Suggestions

Add an explicit validation/feedback loop inside the workflow (e.g., after running garak/PyRIT, parse results, confirm findings reproduce, then iterate) rather than relying solely on the final self-attestation checklist.

For each workflow step, state the concrete success/stop condition (what evidence confirms the step is done) so checkpoints are observable, not just implied by the trailing checkbox list.

Clarify or guard the out-of-bundle relative paths (../ops/skill-supply-chain.md, ../field-journal/precedent-pentest.md, ../tool-index.md) so navigation does not break when those sibling files are absent.

DimensionReasoningScore

Conciseness

The body is information-dense — attack payloads, OWASP mappings, and tool tables — and does not explain basics Claude already knows; every section earns its tokens despite the meta-process framing.

3 / 3

Actionability

Provides copy-paste-ready injection strings across five difficulty levels plus concrete tool install commands ('pip install garak', 'npm install -g promptfoo'), matching the fully-executable anchor.

3 / 3

Workflow Clarity

The six-step workflow is clearly sequenced and ends with a completion checklist, but explicit per-step validation/feedback loops are missing — notable for the batch-probing nature of security testing, which the rubric flags.

2 / 3

Progressive Disclosure

Core methodology is kept inline while detail is delegated to four real one-level-deep reference files (owasp-llm-top10.md, prompt-injection-methodology.md, agent-security-testing.md, agent-obedience-engineering.md), all verified present in the bundle.

3 / 3

Total

11

/

12

Passed

Description

100%Weight 40%Scale 1-3

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A well-crafted description: third-person voice, explicit 'Use for' trigger, and a concrete enumeration of the security domains covered. It reads like the reference good examples and avoids vague fluff.

DimensionReasoningScore

Specificity

Enumerates concrete, specific risk categories — 'prompt injection, tool abuse, RAG exposure, memory poisoning, and model supply-chain risks' — beyond a single vague action, matching the anchor that lists multiple specific concrete items.

3 / 3

Completeness

The 'Use for ...' clause provides explicit trigger guidance answering when, while the enumerated scope answers what, satisfying both halves of the completeness anchor.

3 / 3

Trigger Term Quality

Uses natural terms a user requesting this work would say — 'LLM applications', 'AI agents', 'prompt injection', 'RAG', 'model supply-chain' — giving good coverage of common phrasings.

3 / 3

Distinctiveness Conflict Risk

The LLM/AI-agent security-assessment niche with its specific risk triggers is clearly distinguishable and unlikely to fire for unrelated skills.

3 / 3

Total

12

/

12

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
zhaoxuya520/reverse-skill
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.