CtrlK
BlogDocsLog inGet started
Tessl Logo

review-adversarial

Review a code change for violated assumptions, cross-component composition failures, multi-step failure cascades, abuse through normal use, and verification mechanisms that can pass while production fails. Use when reviewing for adversarial failure scenarios, emergent misbehavior, or green-while-red CI, test, and deploy guards.

74

Quality

93%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

93%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An exemplary instruction-only skill body: dense, concrete, and free of padding, with explicit decision rules for what to report and a concrete output format. The only improvement space is making the review workflow an explicit numbered sequence with per-step checkpoints rather than relying on section order.

DimensionReasoningScore

Conciseness

Lean and efficient with no padding: it never explains concepts Claude already knows (what CI is, what a timeout does) and every clause carries payload — e.g., the assumption examples ('an API always returns JSON, a config key is set, a list is never empty') are the actual working material, not filler. Fits the 'every token earns its place' anchor.

5 / 5

Actionability

For an instruction-only skill the guidance is fully executable: a risk-scaling rule with an explicit high-risk domain list ('authentication, authorization, payments... '), a concrete per-item method ('construct the concrete input or condition, then trace it through the code to its consequence'), explicit report/don't-report decision rules, and a reporting template with a worked title example ('Cascade: payment timeout triggers unbounded retry loop', not 'Missing timeout handling').

5 / 5

Workflow Clarity

The sections mirror a coherent review sequence (scope → risk-scaled method → threshold gate → reporting format) and the Threshold section acts as an explicit validation checkpoint on findings ('Report scenarios you can construct step by step... Do not report speculation'). It falls short of a 5 because the sequence is implicit in section order rather than an explicit ordered workflow with checkpoints for each step.

4 / 5

Progressive Disclosure

The body is ~35 lines, self-contained, with no bundle files present and none needed; under the rubric's simple-skill guidance, well-organized sections with no external references score 5. Scope, Method, Threshold, and Reporting are clearly headed and each appropriately sized for the main file.

5 / 5

Total

19

/

20

Passed

Description

88%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: it enumerates five concrete capability areas and pairs them with an explicit 'Use when...' clause of distinctive trigger phrases. Third-person/imperative voice is correct and there is no fluff. Only gap is modest: a few natural synonyms users might say ('edge cases', 'failure modes') and slight overlap with generic code-review phrasing.

DimensionReasoningScore

Specificity

The description lists multiple specific concrete capabilities — 'violated assumptions, cross-component composition failures, multi-step failure cascades, abuse through normal use, and verification mechanisms that can pass while production fails' — giving comprehensive coverage of the skill's five failure classes, matching the top anchor.

5 / 5

Completeness

It explicitly answers both what ('Review a code change for violated assumptions... verification mechanisms that can pass while production fails') and when ('Use when reviewing for adversarial failure scenarios, emergent misbehavior, or green-while-red CI, test, and deploy guards') with concrete trigger phrases, matching the top anchor exactly.

5 / 5

Trigger Term Quality

Good natural trigger phrases: 'adversarial failure scenarios', 'emergent misbehavior', 'green-while-red CI, test, and deploy guards'. A few common user variations are missing (e.g., 'edge cases', 'failure modes', 'robustness review'), so it fits the anchor of good-but-not-comprehensive keyword coverage rather than the 5's full synonym coverage.

4 / 5

Distinctiveness Conflict Risk

The niche (emergent/adversarial failure review, green-while-red guards) is clearly distinct from neighboring skills, but the opening 'Review a code change' overlaps with general code-review skills, creating minor trigger-conflict risk — fitting 'mostly distinct; minor overlap risk' rather than the 5's minimal-conflict anchor.

4 / 5

Total

18

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
perihelionhq/perihelion-platform-context
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.