CtrlK
BlogDocsLog inGet started
Tessl Logo

critical

Adversarially challenges a proposed plan, code change, or bug diagnosis from a hostile pre-mortem perspective. Walks a fixed taxonomy of failure modes, blast radius, rollback, hidden coupling, and maintainability; every finding must cite a file, line, or named assumption; forces a steelman of at least one alternative. Surfaces concerns only — does not score (delegates to `/confidence`) and does not apply fixes. Use during planning before autonomous execution, before opening a high-stakes PR, or when a fix "feels off". One adversarial pass per run — naïve self-refine loops amplify bias. Modes: plan (default), code, analysis; add `deep` to run 3–5 independent persona lenses in parallel (sub-agents when available, personas found via `ideate`) and merge them — e.g. at the end of a feature. Triggers on "critical", "challenge this", "pre-mortem", "red-team this", "deep critical review", "review from every angle", "/critical".

69

Quality

87%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

77%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a tightly written, checklist-driven adversarial procedure with an explicit output contract, a grounding validation step, and clear mode handling — workflow clarity and actionability for the standard path are excellent. The two real weaknesses are a broken reference (rules/deep-mode.md does not exist in the bundle, leaving deep mode under-specified) and modest rule repetition across sections.

Suggestions

Fix the broken reference: either add the actual rules/deep-mode.md file (with the lens catalog, sub-agent prompt contract, and deep output format the body promises) or inline a minimal fallback lens catalog and sub-agent prompt so deep mode is fully executable from the bundle as shipped.

Deduplicate the repeated hard rules: state the one-pass, every-finding-cites, and no-re-stating rules once (in 'Hard rules and non-goals') and reference them from the persona contract, intro, and output format instead of restating them in each place.

The argument-hint advertises '--lenses <n|a,b-c>' but no section in SKILL.md explains how lens count or named lenses are selected; document that where deep mode is defined.

DimensionReasoningScore

Conciseness

The body is dense and table-driven with essentially no explanation of concepts Claude already knows — the taxonomies, persona blockquote, and output template all earn their tokens. It is not a 5 because several rules are repeated across sections: the single-pass rule appears in the intro, hard rule 1, and the output format; the citation rule appears in both the persona contract and hard rule 4; the no-restatement rule appears as persona rule 3 and hard rule 6. Consolidating these would trim a modest number of tokens without loss. It is not a 3 because beyond this redundancy the content is efficient and assumes the model's competence.

4 / 5

Actionability

The core run is copy-paste executable: the exact persona text to adopt, the mode-announcement line ("Mode: critical/<mode>. Target: <one-line summary>."), the full output-format markdown template, per-mode grounding actions ("Read the plan.md; grep for at least two referenced files/symbols"), and classification definitions for must-fix/should-fix/nice-to-have. It is not a 5 because deep mode defers its "full procedure, lens catalog, sub-agent prompt contract, and output format" to rules/deep-mode.md, and that file does not exist in the bundle — the fallback lens catalog and sub-agent prompt are therefore unavailable, leaving the deep path under-specified.

4 / 5

Workflow Clarity

The sequence is clear and cued explicitly ("State the detected mode in one line before running", grounding "before writing the findings", persona "at the top of every run", ending in the "Next step" section), and validation is explicit: the external grounding rule is a checkpoint whose failure has a defined outcome ("If a referenced file, symbol, or path does not resolve, that is itself a finding (categorised as must-fix)"). The taxonomy tables are checklists (every row walked, skips must be listed), and edge cases have defined behavior ("If Must-fix and Should-fix are both empty, output 'No blocking concerns found.'"). The skill is advisory and read-only, so the destructive/batch validation cap does not apply. It is not a 4 because checkpoints and error-recovery branches are present, not missing.

5 / 5

Progressive Disclosure

Scored against the actual bundle: there are no references/, scripts/, or assets/ directories, and the one path the body points to — [rules/deep-mode.md](./rules/deep-mode.md), cited twice including an anchor (#d1--lens-selection) — does not exist. The in-SKILL.md structure itself is good (a Contents TOC, well-labeled sections, clearly signaled offload), but the promised detail layer (lens catalog, sub-agent prompt contract, deep-mode output format) is neither inline nor present, so a deep-mode run hits a dead end. That is a broken disclosure path rather than a 'minor organization gap', which places it at 3 rather than 4.

3 / 5

Total

16

/

20

Passed

Description

96%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: concrete capabilities, explicit use-when guidance with natural trigger phrases, and deliberate boundary-drawing against sibling skills. The only weakness is that a couple of trigger terms ("critical", "challenge this") are broad enough to compete with other review skills in the same ecosystem.

DimensionReasoningScore

Specificity

The description lists multiple concrete actions with comprehensive coverage: "Adversarially challenges a proposed plan, code change, or bug diagnosis", "Walks a fixed taxonomy of failure modes, blast radius, rollback, hidden coupling, and maintainability", "every finding must cite a file, line, or named assumption", "forces a steelman of at least one alternative", plus explicit non-actions ("does not score (delegates to /confidence) and does not apply fixes"). Third person voice is used throughout ("challenges", "walks", "surfaces"), so no voice penalty applies. It is not a 4 because there are no gaps in action coverage — even the output constraint (one adversarial pass per run) and mode variants (plan/code/analysis, deep) are named.

5 / 5

Completeness

Both 'what' and 'when' are explicit. What: adversarial pre-mortem challenge with taxonomy walk, mandatory citations, mandatory steelman, no scoring, no fixes. When: "Use during planning before autonomous execution, before opening a high-stakes PR, or when a fix 'feels off'" — an explicit 'Use when...' clause, so the completeness cap of 3 does not apply. It is not a 4 because the 'when' is already concrete and trigger-phrased rather than merely present.

5 / 5

Trigger Term Quality

Comprehensive natural-language triggers: "critical", "challenge this", "pre-mortem", "red-team this", "deep critical review", "review from every angle", "/critical", plus scenario phrasings users would actually say ("when a fix 'feels off'", "before opening a high-stakes PR"). These cover formal and casual synonyms of the concept; file extensions do not apply to a process skill. It is not a 4 because the synonym coverage (pre-mortem / red-team / challenge / every angle) spans the realistic phrasing space rather than leaving a few natural terms missing.

5 / 5

Distinctiveness Conflict Risk

The description actively disambiguates from adjacent skills ("does not score (delegates to /confidence) and does not apply fixes", single-pass by design), giving it a clear niche. However, in a review-heavy ecosystem the broad triggers "critical" and "challenge this" carry minor overlap risk with closely related review/quality skills, which matches the 'mostly distinct; minor overlap risk' anchor rather than the 'minimal conflict risk' anchor.

4 / 5

Total

19

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

relative_links

Relative link issues: 2 missing

Warning

Total

13

/

16

Passed

Repository
mthines/agent-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.