CtrlK
BlogDocsLog inGet started
Tessl Logo

llm-risk-assess

Comprehensive LLM security assessment against OWASP Top 10 for LLM Applications 2025. Use when reviewing LLM-integrated applications, RAG pipelines, chatbots, AI agents, or GenAI features. Covers prompt injection, data poisoning, supply chain, excessive agency, and more with real-world attack scenarios and testing methodologies.

62

Quality

74%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./plugins/ai-security-skills/skills/llm-risk-assess/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

61%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a clean, concise overview that delegates detail appropriately, but it stops short of executable guidance and lacks the validation feedback loops expected for a destructive red-team workflow.

Suggestions

Add concrete, runnable commands or probe snippets for the automated testing step (e.g., a sample Garak invocation) to lift actionability.

Insert explicit validation checkpoints in the workflow (e.g., 'confirm probe results before escalating to red-team tests') to satisfy the destructive-operation feedback-loop requirement.

Verify plays/llm-risk-assess.md exists in the bundle, or inline the minimal procedure so the signaled reference is not a dead path.

DimensionReasoningScore

Conciseness

Lean, well-structured bullets that assume Claude's domain knowledge without over-explaining concepts; only minor trimming opportunities exist in the per-category listing.

4 / 5

Actionability

Names specific tools (Garak, Giskard) and concrete attack vectors per OWASP category, but provides no executable code or commands, leaving the guidance high-level relative to the rubric's executable-code anchor.

3 / 5

Workflow Clarity

A clear four-step sequence is present, but validation/verification checkpoints are absent; because red-team testing involves potentially destructive attack execution, the missing feedback loop caps this at 3 per the rubric.

3 / 5

Progressive Disclosure

Good section organization (Steps, Output, OWASP References) with a clearly signaled one-level reference to plays/llm-risk-assess.md for the detailed procedure; minor gaps remain around inline bulk versus delegated detail.

4 / 5

Total

14

/

20

Passed

Description

87%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: it states a concrete purpose, an explicit trigger clause with natural phrasing, and a distinct LLM-security niche. Specificity is the only mild weakness, as threat categories edge out enumerated concrete actions.

DimensionReasoningScore

Specificity

Names the LLM-security domain and lists several concrete threat areas (prompt injection, data poisoning, supply chain, excessive agency) plus testing methodologies, but presents categories more than discrete actions, leaving minor coverage gaps.

4 / 5

Completeness

Explicitly answers both what ('Comprehensive LLM security assessment against OWASP Top 10 for LLM Applications 2025') and when ('Use when reviewing...') with concrete trigger phrases.

5 / 5

Trigger Term Quality

Provides natural trigger phrases users would say ('LLM-integrated applications, RAG pipelines, chatbots, AI agents, or GenAI features') with good coverage, though a few synonyms or phrasings could be added.

4 / 5

Distinctiveness Conflict Risk

Occupies a clear niche (OWASP LLM 2025 security assessment) with distinct triggers tied to LLM/RAG/chatbot/agent contexts, giving minimal conflict risk with other skills.

5 / 5

Total

18

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

Total

15

/

16

Passed

Repository
OWASP/secure-agent-playbook
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.