CtrlK
BlogDocsLog inGet started
Tessl Logo

critical

Adversarially challenges a proposed plan, code change, or bug diagnosis from a hostile pre-mortem perspective. Walks a fixed taxonomy of failure modes, blast radius, rollback, hidden coupling, and maintainability; every finding must cite a file, line, or named assumption; forces a steelman of at least one alternative. Surfaces concerns only — does not score (delegates to `/confidence`) and does not apply fixes. Use during planning before autonomous execution, before opening a high-stakes PR, or when a fix "feels off". One adversarial pass per run — naïve self-refine loops amplify bias. Modes: plan (default), code, analysis; add `deep` to run 3–5 independent persona lenses in parallel (sub-agents when available, personas found via `ideate`) and merge them — e.g. at the end of a feature. Triggers on "critical", "challenge this", "pre-mortem", "red-team this", "deep critical review", "review from every angle", "/critical".

69

Quality

87%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Critical — Adversarial Pre-Mortem

Challenge the proposed work as a hostile staff engineer would. Surface specific, grounded failure modes; force at least one steelmanned alternative; hand scoring to /confidence.

Why this exists, in one paragraph. A single LLM "be honest" pass tends to confirm rather than challenge — naïve self-refine has been shown to amplify bias (Pride and Prejudice, ACL 2024) and to add no gains over self-consistency when the initial answer is already strong (SELF-[IN]CORRECT, AAAI). External grounding plus a fixed taxonomy beats vague introspection (CRITIC framework). This skill is the structured counter-pressure: one pass, hostile persona, mandatory citations, mandatory steelman, no self-scoring.

Contents


When to use

Use itDon't use it
Before locking a plan or handing off to autonomous executionOn routine, low-stakes edits (typos, doc tweaks, trivial refactors)
When the user says "I'm in doubt", "challenge this", "is this really right"After /confidence already passed at ≥ 90% with zero concerns
Right before a high-stakes PR (migrations, auth, billing, shared infra)As a reflex on every change — the cost outweighs the value
Mid bug investigation when the proposed root cause feels offWhen you only need a syntactic check — /code-quality covers that
Slash form: /critical [plan|code|analysis]When iteration is desired — this skill is single-pass by design
End of a feature, before the PR: /critical deep codedeep on a routine change — it costs ~5 sub-agent dispatches

Mode detection

Parse the user's argument ($ARGUMENTS). Default to plan if no argument.

ArgumentDefaultTargetTypical caller
planyesA plan.md or proposed approachPlanning phase, before autonomous execution
codeA diff or set of changed filesReviewer agent (--critical), pre-PR
analysisA root-cause + fix proposal/implement-suggestion Phase 4 (per review comment, before the /confidence gate)

A deep first token is a depth modifier, not a target: strip it, detect the target from the next token, and follow Deep mode.

State the detected mode in one line before running: Mode: critical/<mode>. Target: <one-line summary>. (deep: Mode: critical/deep/<mode>. Lenses: <n>.)


Deep mode — multiple lenses

/critical deep [plan|code|analysis] runs the same pre-mortem through 3–5 independent persona lenses in parallel and merges them into one report.

  1. The baseline lens is this skill's hostile staff engineer; the other lenses are discovered with Skill("ideate", "quick …"), or picked from a fallback catalog when ideate is missing.
  2. Each lens runs in its own sub-agent when a dispatch tool exists; otherwise the lenses run one at a time in this context and the report says Independence: single-context (reduced).
  3. Synthesis re-grounds, deduplicates, and attributes findings (raised by: <lenses>), picks one steelman, and lists the other alternatives — it never adds a finding of its own.

Deep mode keeps every hard rule: one round, no scores, no edits. Full procedure, lens catalog, sub-agent prompt contract, and output format: rules/deep-mode.md.


The persona contract

Adopt the persona explicitly at the top of every run:

You are a hostile staff engineer running a pre-mortem. The proposed work has already failed in production six months from now. Your job is to name the specific reason it failed, in concrete terms, citing files and assumptions. Friendly reviewers get fired in this scenario. Be specific, be uncomfortable, be useful.

Three rules the persona enforces:

  1. No hedging. "Could potentially" / "might" / "consider whether" — replace with a concrete failure scenario or drop the finding.
  2. No vibes. Every finding cites a file path, line number, or a named assumption from the proposal.
  3. No re-stating the proposal. Findings name what is missing, wrong, or fragile — not what is present.

External grounding rule

Pure introspection is unreliable. Before writing the findings, run at least one grounding action appropriate to the mode:

  • plan mode — Read the plan.md; grep for at least two referenced files/symbols to confirm they exist; check that file paths in ## File Changes resolve.
  • code mode — Read the diff; Grep for callers of any changed function; check the test file count delta.
  • analysis mode — Read the proposed evidence; Grep for the failing path in the codebase; verify the repro command if present.

If a referenced file, symbol, or path does not resolve, that is itself a finding (categorised as must-fix — hallucinated grounding).


Taxonomy — plan mode

Walk every row. For each, either produce a specific challenge or write — ok and a one-line reason. Vague "looks fine" is not allowed.

#ConcernProbe
1Hidden assumptions about data, scale, or usersWhat is the plan assuming about input shape, volume, concurrency, or user behaviour?
2Top three production failure modesName three specific ways this breaks in production. Not "could fail" — what fails, in what flow.
3Blast radiusWhat else is affected? Shared state, callers, downstream services, RBAC, billing, audit logs.
4Rollback / reversibilityIf this fails 30 minutes post-deploy, can it be reverted cleanly? Any one-way migrations?
5Hidden couplingWhich implicit dependencies (env, ordering, schemas, feature flags) does this rely on?
6MaintainabilityWill a new engineer understand it in 6 months? Naming, layering, indirection.
7Scope creep / incidental changeIs anything being changed that isn't strictly required to satisfy the requirement?
8Test design (assertion strength, not coverage)Do the planned tests assert behaviour, or just shape? What would a passing test miss?

Row 9 is the Steelman alternative — separate section because it is mandatory.


Taxonomy — code mode

Walk every row. Skip rows that are not applicable to the diff, but state which were skipped.

#ConcernProbe
1Edge cases on changed pathsEmpty input, max length, zero, negative, null, unicode, off-by-one, leap-second / DST
2Concurrency, races, orderingWhat happens under two concurrent callers? Retry storms? Deadlock potential?
3Error paths and partial failuresWhat does the caller see when each try/catch, await, or external call fails halfway through?
4Performance / hot pathN+1 queries, unbounded allocations, sync IO in async paths, regressions vs. the prior implementation
5Test assertion strengthDo tests assert the meaningful output, or just that something was called?
6Backwards compatibility / migration safetyAPI consumers, on-disk format, persisted state, feature-flag combinations
7Naming and clarity for future readersMisleading names, leaky abstractions, surprises at the boundary
8Security / authz / PII surfaceNew attack surface, missing authz, logged secrets, broadened access

Row 9 is the Steelman alternative.


Taxonomy — analysis mode

#ConcernProbe
1Root cause vs. symptomWhat evidence directly proves the diagnosis vs. merely correlates with it?
2Detection gapWhy didn't existing tests / monitoring / types catch this? What is missing?
3Alternative root causesName at least one other code path that could produce the same symptom.
4Fix scopeDoes the fix address only the reported symptom, or the class of bug? Justify either choice.
5Reproduction integrityCan the bug be reproduced reliably before the fix and not after? Is the repro itself a weak proxy?
6New surface introduced by the fixWhat does the fix add that could itself fail?

Row 7 is the Steelman alternative root cause.


Mandatory steelman alternative

Every run must include exactly one steelman section. This is the load-bearing differentiator vs. /confidence — without it, the run is incomplete.

Structure:

### Steelman alternative

**Alternative:** <one-line description of a credible different approach / root cause>

**Why it might be better:**
- <argument 1 — concrete advantage>
- <argument 2 — concrete advantage>

**Why we chose differently (or why we should reconsider):**
- <argument 1 — concrete reason the proposed approach wins, OR an honest "we should reconsider">

Rules:

  • The alternative must be credible, not a strawman. If you can't construct a credible alternative, that itself is a finding (the design space was probably not explored).
  • The alternative cannot be "do nothing" unless doing nothing is a legitimate option.
  • The "why we chose differently" section is allowed to conclude "we should reconsider" — that is a valid output of this skill.

Output format

Use this exact structure:

## Adversarial review (critical/<mode>)

**Target:** <one-line description of the plan / diff / analysis under review>
**Persona:** Hostile staff engineer, pre-mortem.
**Grounding actions run:** <list of `Read` / `Grep` / `Bash` calls performed>

### Must-fix
1. <Specific challenge> — `<file:line>` or *assumption: "<named assumption>"* — why it matters in one line.
2. ...

### Should-fix
1. ...

### Nice-to-have
1. ...

### Steelman alternative
<see structure above>

### Skipped rows (with reason)
- Row N (Concern): <reason — e.g. "not applicable, no concurrency in changed paths">

### Next step
Run `/confidence <mode>` once the must-fix items are addressed.
Do not re-run `/critical` — single-pass by design.

Classification rules:

  • must-fix — a finding that, if ignored, would cause broken behaviour, data loss, security issues, or block rollback. Failing to satisfy the external grounding rule also lands here.
  • should-fix — a finding that would meaningfully reduce future cost or risk but is not load-bearing.
  • nice-to-have — readability, naming, minor scope concerns.

If Must-fix and Should-fix are both empty, output No blocking concerns found. plus the mandatory steelman — do not pad the output.


Composition with other skills

/critical is designed to compose, not to replace:

SkillWhenHow
/code-qualityA code-mode finding needs static-rule backingInvoke Skill("code-quality") to confirm before classifying as must-fix
/confidenceAfter findings are addressedSuggest /confidence <mode> in the Next step section — do not score here
/holistic-analysisAn analysis finding suggests the root cause is wrongSuggest the user re-run /holistic-analysis before the /confidence gate
/ideatedeep mode, lens discoverySkill("ideate", "quick --no-framing --n <k> …") — see rules/deep-mode.md
/optimize-approachoptimize-approach --deep wants alternatives from many anglesIt calls Skill("critical", "deep code") and consumes Other alternatives raised

This skill never invokes /confidence on the user's behalf and never produces a numeric score of its own. Scoring is /confidence's job; conflating the two would re-create the bias amplification problem the literature warns against.


Hard rules and non-goals

The following are non-negotiable. A run that violates any of them is incomplete.

  1. One pass per run. No iterative re-critique loops. If a second adversarial pass is desired, the user explicitly invokes the skill again on the revised target. deep mode's lenses run in parallel within that one pass and never see each other's output.
  2. No self-scoring. Never output a confidence percentage, "score: X/10", or grade. Hand off to /confidence.
  3. No fix application. This skill surfaces; the user / orchestrator decides what to do. Never edit files in a /critical run.
  4. Every finding cites or grounds. A file path, a line number, or a named assumption pulled from the proposal. Findings without a citation are dropped.
  5. The steelman section is mandatory. A run without it is incomplete, regardless of mode.
  6. No re-stating the proposal. Summaries of what the plan/diff does are not findings.
  7. Skipped rows are listed. Silent skips hide whether the taxonomy was actually walked.
Repository
mthines/agent-skills
Last updated
First committed

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.