CtrlK
BlogDocsLog inGet started
Tessl Logo

adversarial-review

Cross-vendor adversarial code review of the current branch. Two different model families (Claude + Codex/GPT) review the diff independently, then try to refute each other's findings; survivors are reported by confidence. Runs from either Claude Code or Codex. Use when the user asks for an adversarial review, a cross-model / second-opinion review, or wants high-confidence findings before merging. Report-only — never auto-applies fixes.

72

Quality

88%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

81%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-engineered orchestration skill: symmetric two-reviewer flow, concrete commands in both directions, mechanical gates, and honest reporting rules with no auto-fix. The main gaps are the partially-specified Phase 2 invocation and bundle files (prompts/, schemas/) that the body depends on but that are not present in this skill directory.

Suggestions

Spell out the Phase 2 refute invocation as a full command block (both Claude and Codex variants), mirroring the Phase 1 treatment, so the whole flow is copy-paste executable.

Include the referenced prompts/review.md, prompts/refute.md, and schemas/findings.schema.json + verdicts.schema.json files in the bundle — the body currently depends on paths that don't exist next to SKILL.md.

Trim incidental asides (e.g. the parenthetical about what 'broke earlier') to push conciseness toward the lean end of the scale.

DimensionReasoningScore

Conciseness

The body is dense and assumes competence — no explanations of what git diffs or CLIs are — with nearly every sentence carrying operational content. It falls just short of the 5 anchor because of trimmable asides like "(the scattered, inconsistent spelling is what broke earlier)" and rationale commentary such as "that independence is the point"; it is clearly above the 3 anchor since there is no padding or teaching of known concepts.

4 / 5

Actionability

Phase 1 gives copy-paste-ready commands with exact flags (`codex exec -s read-only --output-schema ... -o ...` and the `claude -p --permission-mode plan --output-format json` variant) plus the argv-array diff construction. It is not 5 because Phase 2's refute invocation is specified only as a delta ("invoke it again the same way (swap prompts/review.md for prompts/refute.md...)") rather than a full executable command — the 4 anchor ("concrete code or commands with minor gaps") fits.

4 / 5

Workflow Clarity

The sequence is explicit — Inputs, Preflight, Phase 0 gates, Phase 1 independent review, Phase 2 cross-refute, Phase 3 synthesize — with validation checkpoints and error recovery throughout: missing-CLI fallback, empty-diff stop, deterministic gates before the models, JSON output-shape normalization for varying CLI versions, id-matched verdicts, and contested findings never silently dropped. This matches the 5 anchor (clear sequence, explicit validation, feedback loops); it is not 4 because no checkpoint is merely implicit.

5 / 5

Progressive Disclosure

Structure is good: a ~150-line orchestration overview that appropriately pushes the review brief, refute brief, and output shapes to one-level-deep bundle paths ("prompts/review.md", "schemas/findings.schema.json") resolved relative to the skill dir. It is not 5 because the referenced prompts/ and schemas/ files are not present in the bundle alongside SKILL.md, so the navigation targets cannot actually be reached from what ships here — the 4 anchor ("references mostly clear; minor organization gaps").

4 / 5

Total

17

/

20

Passed

Description

96%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An excellent description: concrete third-person actions, explicit use-when triggers with natural synonyms, and clear scope boundaries (report-only, both host CLIs). The only weak point is a minor trigger overlap with ordinary code-review skills.

Suggestions

Sharpen the when-clause to exclude routine reviews, e.g. 'Use when the user explicitly asks for an adversarial or cross-model review — not for ordinary single-model review requests.'

DimensionReasoningScore

Specificity

The description lists multiple concrete actions — "Two different model families (Claude + Codex/GPT) review the diff independently, then try to refute each other's findings; survivors are reported by confidence" — plus explicit constraints ("Runs from either Claude Code or Codex", "Report-only — never auto-applies fixes"), all in third-person voice. This matches the 5 anchor (multiple specific concrete actions, comprehensive coverage); it is not the 4 anchor because there are no meaningful gaps in coverage of what the skill does.

5 / 5

Completeness

Both questions are answered explicitly: what — "Cross-vendor adversarial code review of the current branch... survivors are reported by confidence" — and when — "Use when the user asks for an adversarial review, a cross-model / second-opinion review, or wants high-confidence findings before merging." This matches the 5 anchor exactly (clear what AND when with concrete trigger phrases); the 4 anchor ("when could be more explicit or specific") does not apply since the when-clause enumerates concrete user phrasings.

5 / 5

Trigger Term Quality

It covers the natural phrases a user would actually say with synonyms included: "adversarial review", "a cross-model / second-opinion review", "high-confidence findings before merging". This matches the 5 anchor (comprehensive coverage including synonyms); it is not 4 because the trigger list already spans the natural synonym space (adversarial / cross-model / second-opinion) rather than missing common variants.

5 / 5

Distinctiveness Conflict Risk

The niche is clear — cross-vendor adversarial review with refutation — with distinct trigger words ("adversarial", "cross-model", "second-opinion"), but the opening "code review of the current branch" overlaps with generic code-review skills for plain "review my branch" requests. This matches the 4 anchor (mostly distinct, minor overlap risk with closely related skills); it is not 5 because a routine review request could plausibly match both this and a standard review skill.

4 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
basicmachines-co/basic-memory
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.