CtrlK
BlogDocsLog inGet started
Tessl Logo

autoreview

Structured Codex, Claude, Amp, Pi, or Kimi code review when explicitly requested.

56

Quality

63%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.agents/skills/autoreview/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

57%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A thorough, highly actionable reference for running a multi-engine structured review, with concrete commands and validation built into the bundle workflow. Its main weaknesses are density/verbosity in the isolation and provenance prose and a 427-line monolithic body that inlines reference material instead of splitting it into separate files.

Suggestions

Split the long 'Review engine isolation' paragraph (line 332) and the model/env tables into a separate reference file (e.g., references/ENGINES.md) and link to it, shrinking SKILL.md toward an overview.

Tighten run-on sentences in the Contract and Oversized Bundles sections into shorter bullet clauses to improve token efficiency.

Add an explicit numbered workflow (set paths → pick target → run → handle results → report) with validation checkpoints called out, so the sequence is visible rather than implied by section order.

DimensionReasoningScore

Conciseness

Most content is specialized engine/isolation detail Claude would not already know, but several sections are dense run-on prose (e.g., the 332-line 'Review engine isolation' paragraph and the line-36 TruffleHog sentence) that could be tightened without loss, fitting 'mostly efficient but includes some unnecessary explanation or could be tightened'.

3 / 5

Actionability

It provides many copy-paste-ready bash commands across local/branch/commit modes plus concrete model, flag, and environment-variable tables, but the Contract and Scope sections are policy-level prose rather than executable guidance, leaving minor gaps.

4 / 5

Workflow Clarity

Validation checkpoints exist (freeze-and-scan before each provider call, structured validation, exit codes) but the content is organized as a topical reference manual rather than a sequenced workflow, so checkpoints are scattered and implicit rather than a clear numbered sequence with feedback loops.

3 / 5

Progressive Disclosure

Section headers are clear and references are one level deep (scripts via $AUTOREVIEW/$AUTOREVIEW_HARNESS, --help, external doc URLs), but the 427-line SKILL.md inlines substantial reference-grade tables (isolation flags, model defaults, env vars) with no reference/ files to split them into, fitting 'content that should be separate is inline'.

3 / 5

Total

13

/

20

Passed

Description

70%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A concise, third-person description that clearly states what the skill does and gates it on explicit request, with good engine-name keyword coverage. It could name the concrete trigger phrases ('autoreview', 'second-model review') and a second concrete action to reach the top anchors.

Suggestions

Add the natural trigger phrase 'autoreview' (and 'second-model review') to the description so it matches how users actually invoke the skill.

Expand 'when explicitly requested' into a concrete 'Use when the user asks for autoreview, a second-model review, or a named engine review' clause.

Add one more concrete action (e.g., '...and reports blocking P0 findings') to lift specificity toward the several-actions anchor.

DimensionReasoningScore

Specificity

The description names the domain ('code review') and five concrete engines ('Codex, Claude, Amp, Pi, or Kimi') plus a 'Structured' qualifier, but offers only a single action rather than multiple concrete actions, fitting the anchor that names the domain with limited actions better than the several-actions anchor.

3 / 5

Completeness

It answers both what ('Structured ... code review') and when ('when explicitly requested'), satisfying the explicit-trigger requirement above the 3 cap, but the 'when' is generic rather than naming concrete trigger phrases, matching the anchor where 'when' could be more specific.

4 / 5

Trigger Term Quality

It includes natural terms users would say ('code review', 'Codex', 'Claude', 'Amp', 'Pi', 'Kimi') with good coverage, but omits natural trigger variants like the skill's own name 'autoreview' and 'second-model review' that the body itself relies on, so it stops short of comprehensive.

4 / 5

Distinctiveness Conflict Risk

The multi-engine structured-review niche plus 'when explicitly requested' makes it mostly distinct with only minor overlap risk against generic code-review skills or Guardian's auto_review, fitting the 'mostly distinct; minor overlap risk' anchor rather than the minimal-conflict anchor.

4 / 5

Total

15

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
openclaw/crabbox
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.