CtrlK
BlogDocsLog inGet started
Tessl Logo

aatmf-t11-agentic-exploit

AATMF T11 — Agentic & Orchestrator Exploitation. MCP tool poisoning, agent-to-agent prompt injection, tool-result spoofing, orchestrator state confusion.

53

Quality

60%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Critical

Do not install without reviewing

Fix and improve this skill with Tessl

tessl review fix ./packages/decepticon/decepticon/skills/plugins/llm-redteam/t11-agentic-exploit/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

61%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a well-structured, lean technique taxonomy with a concrete probe-pattern config, but it lacks an explicit multi-step testing workflow with validation checkpoints and keeps methodology guidance at a high level.

Suggestions

Turn the testing approach into an explicit numbered sequence (build malicious MCP → register → run probe plugins → inspect detection signals) with a validation/check step to confirm whether an exploit succeeded.

Expand the probe pattern into actionable commands or a minimal script for building and registering the malicious test MCP server, rather than leaving it as a one-line instruction.

Move the per-technique detail (T11.001–T11.006) into a reference file and keep SKILL.md as an overview pointer, improving progressive disclosure.

DimensionReasoningScore

Conciseness

Terse bullet style and abbreviations ("w/", "exfils") keep it lean and assume Claude's competence; only minor fluff like the "biggest emerging attack class" claim could be trimmed.

4 / 5

Actionability

The probe-pattern YAML is concrete and copy-pasteable, but surrounding methodology ("build a malicious test MCP server + register it... observe behavior") is high-level guidance rather than executable steps.

3 / 5

Workflow Clarity

An implicit flow (probe → detect → severity → defend) is present via section ordering, but there is no explicit step sequence or validation checkpoints for the testing process.

3 / 5

Progressive Disclosure

Well-organized with clear section headers and no nested references; all content is inline and self-contained, though at ~107 lines some technique detail could be split into reference files.

4 / 5

Total

14

/

20

Passed

Description

58%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description clearly names the agentic-exploitation niche and lists concrete technique categories, but it is missing an explicit 'Use when...' trigger clause and leans on technical jargon over natural user phrasing.

Suggestions

Add an explicit 'Use when...' clause naming natural trigger phrases (e.g., 'Use when testing MCP servers, multi-agent systems, or tool-calling LLMs for injection and tool-abuse vulnerabilities').

Soften jargon with synonyms a user might actually say (e.g., 'agent security', 'tool abuse', 'LLM tool-call attacks') alongside the AATMF labels.

Consider naming all major sub-techniques or summarizing coverage so the description matches the body's full scope.

DimensionReasoningScore

Specificity

Lists several concrete technique categories ("MCP tool poisoning, agent-to-agent prompt injection, tool-result spoofing, orchestrator state confusion") but omits two of the six body techniques, leaving minor coverage gaps.

4 / 5

Completeness

Has a clear 'what' (the technique catalog) but no 'Use when...' clause or equivalent trigger guidance, which caps completeness at 3 per the rubric.

3 / 5

Trigger Term Quality

Contains relevant domain keywords ("MCP tool poisoning", "prompt injection", "tool-result spoofing") but they skew toward technical jargon and lack natural user-spoken synonyms or variations.

3 / 5

Distinctiveness Conflict Risk

The "AATMF T11 — Agentic & Orchestrator Exploitation" niche is distinct, but overlaps with related tactics (T1 prompt injection, T13 supply chain) referenced in the body, leaving minor conflict risk.

4 / 5

Total

14

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

Total

15

/

16

Passed

Repository
PurpleAILAB/Decepticon
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.