CtrlK
BlogDocsLog inGet started
Tessl Logo

aatmf-t03-reasoning-exploit

AATMF T3 — Reasoning & Constraint Exploitation. System prompt override, constraint negation, role-reversal, instruction conflict exploit.

52

Quality

58%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Critical

Do not install without reviewing

Fix and improve this skill with Tessl

tessl review fix ./packages/decepticon/decepticon/skills/plugins/llm-redteam/t03-reasoning-exploit/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

65%

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The content is a lean, well-structured red-team technique catalog that assumes Claude's competence and provides a concrete probe pattern plus detection and mitigation guidance. Its main gaps are the lack of an executable/sequenced workflow with validation checkpoints and the absence of progressive split-out of deeper material.

Suggestions

Add a short sequenced workflow for running a T3 probe (e.g., select technique -> configure probe YAML -> run -> inspect detection signals -> score severity -> apply defender checklist) with an explicit validation/verification checkpoint.

Make the technique entries more actionable by pairing each illustrative attack phrase with the concrete probe input or prompt template to use, rather than relying on quoted example strings.

Consider splitting the per-technique detail (T3.001-T3.006) into a referenced reference file and keeping SKILL.md as an overview, to improve progressive disclosure for the longer catalog.

DimensionReasoningScore

Conciseness

The body is terse and dense: short technique capsules, a compact probe-pattern block, and lean detection/severity/defender lists with no explanations of concepts Claude already knows. It matches the score-3 anchor (lean, every token earns its place) and is not score 2 because there is no padding to trim.

3 / 3

Actionability

The probe-pattern YAML is concrete and config-ready, and detection/defender sections give usable lists, but the core technique entries are illustrative attack phrases rather than executable operational steps. Per the code-vs-instruction note absence of code is acceptable, yet the guidance is still partly descriptive, so it sits at score 2 rather than 3 and above score 1 because concrete probe config and mitigations are present.

2 / 3

Workflow Clarity

The document is well-organized into labeled sections but is a taxonomy/reference, not a sequenced multi-step workflow, and there are no validation checkpoints or feedback loops. It is not score 3 because no explicit sequence with validation exists, and not score 1 because the sections are clearly structured and ordered.

2 / 3

Progressive Disclosure

It is a single ~94-line file with clean section headers and only external tactic cross-references (T1/T10/T11) rather than nested file references, but all content is inline with no progressive split into referenced materials. It is not score 3 because nothing is split out for deeper reading, and not score 1 because organization is clear and references are shallow.

2 / 3

Total

9

/

12

Passed

Description

52%

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific and clearly scoped to a distinct AI-security niche, but it relies on technical jargon rather than natural trigger terms and omits any explicit "when to use" guidance. Adding a "Use when..." clause with user-friendly trigger phrases would most improve it.

Suggestions

Add an explicit 'Use when...' clause stating when Claude should invoke this skill (e.g., 'Use when red-teaming LLM reasoning vulnerabilities, jailbreak testing, or probing policy adherence via indirect prompts').

Replace jargon-only terms with natural-language trigger variations users would actually say, such as 'jailbreak', 'prompt injection testing', 'bypassing safety rules', and 'getting the model to reveal its instructions'.

Keep the concrete technique list but pair each with a plain-language equivalent so the description works both as a capability statement and a discovery trigger.

DimensionReasoningScore

Specificity

"System prompt override, constraint negation, role-reversal, instruction conflict exploit" lists multiple specific concrete techniques, matching the score-3 anchor. It is not score 2 because it enumerates several distinct actions rather than naming a domain with only partial actions.

3 / 3

Completeness

The description answers "what" (lists the exploit techniques) but lacks any "Use when..." or equivalent explicit trigger guidance for when Claude should invoke it. Per the guidelines a missing trigger clause caps completeness at 2; it is not 3 because "when" is only implied, and not 1 because "what" is clearly stated.

2 / 3

Trigger Term Quality

Terms like "constraint negation", "role-reversal", and "instruction conflict exploit" are red-team/attacker jargon rather than phrases a user would naturally say. It is not score 2 because none of the keywords are common natural-language variations a user would utter.

1 / 3

Distinctiveness Conflict Risk

AATMF T3 occupies a clear niche (AI-security reasoning-exploit red-teaming) with distinct triggers unlikely to overlap with general skills. It is not score 2 because the scope is sharply bounded to a specific attack tactic.

3 / 3

Total

9

/

12

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

Total

15

/

16

Passed

Repository
PurpleAILAB/Decepticon
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.