CtrlK
BlogDocsLog inGet started
Tessl Logo

aatmf-t03-reasoning-exploit

AATMF T3 — Reasoning & Constraint Exploitation. System prompt override, constraint negation, role-reversal, instruction conflict exploit.

56

Quality

63%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Critical

Do not install without reviewing

Fix and improve this skill with Tessl

tessl review fix ./packages/decepticon/decepticon/skills/plugins/llm-redteam/t03-reasoning-exploit/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

68%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a concise, actionable red-team playbook with concrete attack exemplars and a usable probe config, organized into clear sections. Its main weakness is the absence of an explicit sequenced workflow with validation checkpoints for the testing process it implies.

Suggestions

Add an explicit end-to-end workflow (e.g., 1. Run probe pattern → 2. Collect outputs → 3. Check detection signals → 4. Score severity) with a validation/retry checkpoint after detection.

Note which plugin ids in the Probe pattern are illustrative vs. real harness plugins, or point to where they are defined.

Expand the telegraphic shorthand ('w/o', bare arrows) into minimal complete sentences in the technique intros so guidance is unambiguous without inference.

DimensionReasoningScore

Conciseness

The body is lean and assumes Claude's competence — terse technique descriptions, arrow notation, and 'w/o' shorthand with no padding or basic-concept explanations — though the telegraphic style occasionally verges on cryptic rather than clean.

4 / 5

Actionability

Each technique ships concrete example attack strings and a copy-paste YAML 'Probe pattern' with plugin ids, numTests, and strategies; mostly executable guidance with minor gaps since plugin ids like 'system-prompt-override' are illustrative rather than verified-real.

4 / 5

Workflow Clarity

The implied probe → detection → severity flow is present via section ordering, but there is no explicit sequenced workflow with validation checkpoints — acceptable for a reference catalog, but the testing process lacks checkpoints and feedback loops.

3 / 5

Progressive Disclosure

Well-organized into clearly signaled sections (Techniques, Probe pattern, Detection signals, Severity, Defender, Cross-references) with one-level cross-references to T1/T10/T11; no bundle files exist but the single-file structure is appropriately sectioned for its size.

4 / 5

Total

15

/

20

Passed

Description

58%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is a third-person capability catalog with concrete, distinct technique names, but it reads as a taxonomy label rather than an invokable skill description and lacks an explicit 'Use when' trigger clause in the description field itself.

Suggestions

Add an explicit 'Use when...' trigger clause to the description (e.g., 'Use when red-teaming LLM reasoning chains, jailbreaking via constraint negation or role-reversal, or testing system-prompt extraction').

Drop or de-emphasize the opaque 'AATMF T3 —' code identifier from the user-facing description so the leading phrase is a natural trigger term.

Fold the metadata.when_to_use keywords into the description sentence so trigger terms appear where Claude actually reads them.

DimensionReasoningScore

Specificity

Lists several specific concrete actions — 'System prompt override, constraint negation, role-reversal, instruction conflict exploit' — naming the domain plus multiple techniques, with only minor coverage gaps (no mention of stepwise collapse or extraction).

4 / 5

Completeness

Gives a clear 'what' (a catalog of reasoning/constraint exploitation techniques) but the description field has no explicit 'Use when...' trigger clause — the when-guidance lives only in metadata.when_to_use, which per the rubric caps completeness at 3.

3 / 5

Trigger Term Quality

Contains relevant security-testing keywords ('system prompt override', 'constraint negation', 'role-reversal') but the leading 'AATMF T3 — Reasoning & Constraint Exploitation' is opaque taxonomy jargon, and plain-language variations users would naturally say are sparse.

3 / 5

Distinctiveness Conflict Risk

The niche is well-scoped to reasoning/constraint-exploitation tactics and the unique 'AATMF T3' identifier plus technique-specific triggers give it low overlap risk, though it sits within a broader AI-security family that creates minor related-skill overlap.

4 / 5

Total

14

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

Total

15

/

16

Passed

Repository
PurpleAILAB/Decepticon
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.