CtrlK
BlogDocsLog inGet started
Tessl Logo

llm-redteam-overview

LLM red team category — full AATMF v3 tactic coverage (T01–T15). Routing skill: read this first to identify which tactic applies, then load the matching sub-skill. Maps to MITRE ATLAS where overlap exists.

60

Quality

71%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Critical

Do not install without reviewing

Fix and improve this skill with Tessl

tessl review fix ./packages/decepticon/decepticon/skills/plugins/llm-redteam/SKILL.md
SKILL.md
Quality
Evals
Security

LLM Red Team Skill Catalog — AATMF v3

This is the routing skill for AI/LLM red-team work. The 15 sub-skills below cover every tactic in the AI/ML Adversarial Tactics, Techniques & Mitigations Framework (AATMF v3). Load the specific sub-skill that matches the attacker objective.

Tactic Map

TacticSub-SkillCoversLoad Path
T01prompt-injectionDirect + indirect prompt injection, ASCII smuggling, payload-in-image, prompt leakingload_skill("/skills/plugins/llm-redteam/t01-prompt-injection/SKILL.md")
T02linguistic-evasionTranslation/transliteration bypass, base64/leetspeak/emoji encoding, low-resource-language jailbreak, multi-lingual context splitload_skill("/skills/plugins/llm-redteam/t02-linguistic-evasion/SKILL.md")
T03reasoning-exploitCoT / ReAct hijack, math/logic distractor, role-play escalation, hypothetical / counterfactual framingload_skill("/skills/plugins/llm-redteam/t03-reasoning-exploit/SKILL.md")
T04memory-manipulationLong-context overflow, conversation rewrite, sliding-window poisoning, "previous turn" forgeryload_skill("/skills/plugins/llm-redteam/t04-memory-manipulation/SKILL.md")
T05api-exploitationFunction-calling abuse, tool-schema confusion, parameter pollution, response-format coercionload_skill("/skills/plugins/llm-redteam/t05-api-exploitation/SKILL.md")
T06training-poisoningBackdoor trigger injection, label flip, RLHF reward hacking, fine-tune dataset contaminationload_skill("/skills/plugins/llm-redteam/t06-training-poisoning/SKILL.md")
T07output-exfilData-leak via reflection, side-channel via length/timing, watermark stripping, token-by-token exfilload_skill("/skills/plugins/llm-redteam/t07-output-exfil/SKILL.md")
T08deceptionConfident-hallucination weaponization, persona impersonation, source spoofing, "as the system says" framingsload_skill("/skills/plugins/llm-redteam/t08-deception/SKILL.md")
T09multimodalImage/audio prompt injection, OCR-payload, steganographic prompts, adversarial perturbationsload_skill("/skills/plugins/llm-redteam/t09-multimodal/SKILL.md")
T10confidentiality-breachSystem-prompt extraction, weight inference, training-data extraction, PII echoload_skill("/skills/plugins/llm-redteam/t10-confidentiality-breach/SKILL.md")
T11agentic-exploitTool-chain hijack, autonomous-loop poisoning, plan-injection, sub-agent confusionload_skill("/skills/plugins/llm-redteam/t11-agentic-exploit/SKILL.md")
T12rag-poisoningRAG index injection, document smuggling, embedding-collision, retrieval-rank gamingload_skill("/skills/plugins/llm-redteam/t12-rag-poisoning/SKILL.md")
T13supply-chainModel registry tampering, dependency confusion (HF Hub, Ollama, MCP), serialization payload abuseload_skill("/skills/plugins/llm-redteam/t13-supply-chain/SKILL.md")
T14infra-warfareGPU resource abuse, billing-amplification, rate-limit DoS, cold-start abuse, region failover poisoningload_skill("/skills/plugins/llm-redteam/t14-infra-warfare/SKILL.md")
T15human-ai-couplingOperator manipulation via AI output, social engineering uplift, dark-pattern UX exploiting AI trustload_skill("/skills/plugins/llm-redteam/t15-human-ai-coupling/SKILL.md")

Quick Routing

Target type identified?
├── Chat-only LLM endpoint     → T01, T02, T03, T10
├── LLM with tools / function calling → T05, T11
├── LLM with RAG / retrieval     → T12, T01 (indirect)
├── LLM with multimodal input    → T09
├── Fine-tuned / custom model    → T06, T10 (extraction)
├── Cross-tenant / shared infra  → T14, T13
└── Human-in-the-loop product    → T15, T08

Tooling

ToolUse
promptfooPlugin-based red-team eval — declarative test cases, runs all 15 tactic plugins
garakNVIDIA's LLM vulnerability scanner — alignment, jailbreak, leak probes
pyritMicrosoft's open red-team automation framework
agentdojoAgentic-system specific (T11) — sandboxed tool-chain attack benchmark
LangSmith / LangfuseObservation surface — replay attacks against deployed agents

Cross-Reference

  • MITRE ATLAS — covers tactic-level coverage (Reconnaissance → Impact). AATMF is more attacker-procedure focused. Most AATMF techniques map to one or more ATLAS techniques; the sub-skills note the mapping in their frontmatter.
  • OWASP LLM Top 10 — risk-class level. AATMF tactics map across LLM01-LLM10.
  • NIST AI RMF AI 600-1 — governance lens; AATMF gives the offensive counterpart.

When in doubt

If the target's surface or behavior doesn't cleanly match one tactic, run a T01 + T05 probe first — direct prompt injection and function-call abuse are the most common, highest-yield starting points and their results inform which other tactics are reachable.

Repository
PurpleAILAB/Decepticon
Last updated
First committed

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.