CtrlK
BlogDocsLog inGet started
Tessl Logo

adversarial-document-reviewer

Conditional document-review persona, selected when the document has >5 requirements or implementation units, makes significant architectural decisions, covers high-stakes domains, or proposes new abstractions. Challenges premises, surfaces unstated assumptions, and stress-tests decisions rather than evaluating document quality.

64

Quality

79%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.agents/skills/adversarial-document-reviewer/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

75%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, lean persona skill: concrete adversarial techniques with specific probes, an explicit depth-calibration gate, confidence thresholds that serve as validation checkpoints, and a territory section that avoids overlap with sibling reviewers. Uniformly strong with only minor gaps — a vague 'Standard' depth tier and no findings-output format.

Suggestions

Quantify the 'Standard' depth tier (e.g., '1000-3000 words or 5-10 requirements, one or two risk signals') so the middle case is as decidable as Quick and Deep.

Add brief guidance on the findings output format (e.g., one block per finding: premise/assumption quoted, counterargument, consequence, confidence) so results are consistently structured.

State explicitly in the depth-calibration section which techniques map to which tier when signals conflict (e.g., few requirements but a high-stakes domain), since the three tier definitions can overlap.

DimensionReasoningScore

Conciseness

Nearly every bullet instructs a technique with concrete conditions ("what happens at 10x? At 0.1x?", "High reversal cost + low evidence quality = risky decision") and nothing explains concepts Claude already knows. It sits just below the lean-everywhere anchor 5 because the depth-calibration tier descriptions and a few bullets could be tightened slightly.

4 / 5

Actionability

For an instruction-only skill, the guidance is highly actionable: each of the five techniques has named probes with specific conditions and consequences (falsification test, reversal cost, subtraction test, do-nothing baseline). The minor gap keeping it from 5 is that the 'Standard' depth tier is defined vaguely ("medium document, moderate complexity") while Quick and Deep have quantified thresholds.

4 / 5

Workflow Clarity

The sequence is clear: calibrate depth from explicit size/risk signals, run the technique set for that depth, then apply confidence calibration, with "Below 0.50: Suppress" acting as an explicit validation checkpoint. Anchor 4 fits: most checkpoints are present, but there is no explicit guidance on how to structure or order the findings output itself.

4 / 5

Progressive Disclosure

The skill is a single self-contained file with well-organized, clearly headed sections (depth calibration, five techniques, confidence calibration, territory) and no bundle files to reference; nothing clearly belongs in a separate file. The under-50-lines exception for a score of 5 does not apply at this length, so the good-structure anchor 4 is the best fit.

4 / 5

Total

16

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, third-person description that clearly states both what the skill does and the quantified conditions under which it is selected. Trigger terms are natural and specific, with only minor gaps in synonym coverage and a modest overlap risk inherent to the document-review domain.

DimensionReasoningScore

Specificity

The description lists several concrete actions ("Challenges premises, surfaces unstated assumptions, and stress-tests decisions") plus explicit selection conditions, matching the 'several specific actions; minor gaps' anchor. It falls short of 5 because it does not enumerate the fuller technique set (e.g., simplification pressure, alternative blindness).

4 / 5

Completeness

It explicitly answers both questions: what ("Challenges premises, surfaces unstated assumptions, and stress-tests decisions") and when ("selected when the document has >5 requirements or implementation units, makes significant architectural decisions, covers high-stakes domains, or proposes new abstractions") with concrete trigger conditions. The 'when' is specific and quantified, so the anchor-4 weaker-'when' case does not apply.

5 / 5

Trigger Term Quality

Good keyword coverage with natural terms a user would say: "requirements", "architectural decisions", "high-stakes", "abstractions", "premises", "assumptions". A few natural phrasings are missing (e.g., "design review", "plan critique", "spec"), which keeps it below the comprehensive-synonyms anchor 5.

4 / 5

Distinctiveness Conflict Risk

The description carves out a distinct epistemological niche and explicitly differentiates ("rather than evaluating document quality"), giving mostly distinct triggers. It stays at 4 rather than 5 because it operates in the crowded document-review space where overlap with other review skills remains a minor risk.

4 / 5

Total

17

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

13

/

16

Passed

Repository
udecode/plate
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.