CtrlK
BlogDocsLog inGet started
Tessl Logo

ai-llm-app-attack

AI/LLM应用攻击:提示注入,Agent工具滥用RCE,RAG投毒,MCP供应链,torch.load pickle RCE。Use when testing LLM apps, agents, RAG, MCP plugins, or AI model file risks.

64

Quality

78%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Critical

Do not install without reviewing

Fix and improve this skill with Tessl

tessl review fix ./skills/ai-llm-app-attack/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

68%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An extremely token-efficient AI/LLM attack-surface cheat sheet that adds genuine non-obvious knowledge, but it reads as a taxonomy rather than an operational playbook — no ordered workflow, commands, or payloads. Its one verification rule (side-effect confirmation before recording a Fact) is a strong touch that deserves expansion.

Suggestions

Convert the taxonomy into a short ordered test workflow (e.g., 1. discover endpoints/tools → 2. probe indirect-injection vectors in RAG/tool outputs → 3. verify side effects via OOB callback → 4. only then record a Fact), so workflow_clarity reaches the explicit-sequence bar.

Add 2-3 copy-paste-ready concrete techniques — e.g., a sample cloud-metadata SSRF URL for the fetch tool, a minimal indirect prompt-injection payload template, and a safe torch.load alternative or detection command — to move actionability from descriptive to executable.

Split the single dense code block into three short headed sections (Attack vectors / Verification / Endpoint discovery) so the cheat sheet scans faster without adding tokens.

DimensionReasoningScore

Conciseness

The body is an 8-line dense cheat sheet with zero padding and no explanation of concepts Claude already knows — every line carries new attack-surface facts. It matches the 'lean and efficient; every token earns its place' anchor exactly.

5 / 5

Actionability

There is concrete guidance ("fetch工具→SSRF内网/云元数据", "文件工具→读/etc/passwd写webshell", "发现端点:抓流量找/chat /agent /tool,问Agent'你有哪些工具'") but no executable commands, payloads, or specific test procedures — it is a taxonomy of attack categories rather than instructive steps, fitting the 'some concrete guidance but incomplete; missing key details' anchor. It is not 4 because the anchor requires concrete code or commands with only minor gaps, and none are present.

3 / 5

Workflow Clarity

One explicit validation checkpoint exists ("验证:实际触发工具副作用(OOB回连/读到文件)才写Fact") and a loose discovery-to-exploitation-to-verification flow is implied, but there is no explicit sequenced workflow — content is a knowledge map, matching 'steps listed but checkpoints missing or implicit'. It is not 2 because validation is explicitly stated rather than absent, and not 4 because no ordered procedure is given.

3 / 5

Progressive Disclosure

The skill is under 50 lines with no bundle files (no references/, scripts/, or assets/ exist), and the single heading plus compact block is reasonably well organized — but everything is crammed into one dense code block that could be split into clearly headed sections (vectors / verification / discovery). This fits 'good structure; minor organization gaps' rather than the fully well-organized-sections bar for a 5.

4 / 5

Total

15

/

20

Passed

Description

88%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A terse but highly specific description that clearly states concrete AI/LLM attack capabilities and pairs them with an explicit 'Use when' trigger clause. Main weakness is slightly limited natural-language synonym coverage for triggering contexts.

DimensionReasoningScore

Specificity

The description lists multiple concrete attack capabilities — "提示注入" (prompt injection), "Agent工具滥用RCE", "RAG投毒", "MCP供应链", "torch.load pickle RCE" — each a specific, verifiable action rather than vague language, with comprehensive coverage of the AI/LLM attack surface. It clearly exceeds the 'several specific actions; minor gaps' anchor (4) and matches the 'multiple specific concrete actions; comprehensive coverage' anchor.

5 / 5

Completeness

It explicitly answers both 'what' (attack categories: prompt injection, agent tool abuse to RCE, RAG poisoning, MCP supply chain, torch.load pickle RCE) and 'when' ("Use when testing LLM apps, agents, RAG, MCP plugins, or AI model file risks") with concrete trigger phrases, matching the top anchor. It is not score 4 because the 'when' clause is already explicit and specific rather than needing more detail.

5 / 5

Trigger Term Quality

Trigger phrases "Use when testing LLM apps, agents, RAG, MCP plugins, or AI model file risks" give good natural-term coverage (LLM apps, agents, RAG, MCP), but common synonyms a user might say — 'prompt injection', 'red team', 'AI security', 'chatbot' — are missing (some appear only in Chinese in metadata tags). This fits the 'good keyword coverage; a few natural terms missing' anchor and falls short of the comprehensive-with-synonyms anchor (5).

4 / 5

Distinctiveness Conflict Risk

The AI/LLM-app-attack niche with distinct triggers (LLM, agents, RAG, MCP plugins, model files) is well-differentiated, matching 'mostly distinct; minor overlap risk'. It is not 5 because the terms could still overlap with a general penetration-testing or security-review skill, and not 3 because RAG/MCP/LLM-app triggers are strongly specific to this niche.

4 / 5

Total

18

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

metadata_field

'metadata' should map string keys to string values

Warning

Total

14

/

16

Passed

Repository
AIPentest/CyberStrikeAI
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.