CtrlK
BlogDocsLog inGet started
Tessl Logo

autoreview

Pre-commit/ship code review: Codex default; optional Claude or Pi.

51

Quality

64%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.agents/skills/autoreview/SKILL.md

The canonical home for this skill is autoreview in openclaw/openclaw

SKILL.md
Quality
Evals
Security

Quality

Content

63%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is highly actionable with concrete, copy-paste-ready commands, real validation feedback loops, and correctly externalized executable scripts. Its main weaknesses are verbosity — exhaustive edge-case policy inlined as dense prose — and underuse of progressive disclosure, with model/isolation/env-var reference material that belongs in separate reference files living inside SKILL.md.

Suggestions

Move the engine-isolation tables, model/thinking tables, and environment-default tables into a references/ file (e.g. references/engines.md), keeping SKILL.md to the core workflow with one-line pointers.

Compress the long single-paragraph contract bullets (TruffleHog policy, automerge provenance, Testbox/TMPDIR details) into short normative rules; push edge-case minutiae to a reference file.

Reorder sections to follow the actual execution sequence (paths → target → run → findings → report) so the workflow reads linearly instead of interleaving policy sections with operational ones.

DimensionReasoningScore

Conciseness

The ~440-line body is noticeably verbose: long single-paragraph contract bullets (e.g. the TruffleHog/secret-scanning paragraph, clawsweeper automerge provenance rules, TMPDIR/Unix-socket detours) pack exhaustive edge cases inline that could each be one short line. It avoids explaining concepts Claude already knows, so it is not a 1 ('severely verbose... padded with known concepts'), but it sits well below the 'mostly efficient' midpoint of 3 given the sheer volume of conditional policy detail.

2 / 5

Actionability

The body is fully executable: copy-paste-ready commands for every common case ("$AUTOREVIEW" --mode local / --mode branch --base origin/main / --mode commit --commit HEAD, gh pr view --json baseRefName --jq .baseRefName, --reviewers codex,claude,pi), plus concrete flag tables, env-var defaults, and a runnable smoke harness ("$AUTOREVIEW_HARNESS" --fixture benign --engine codex) whose referenced files exist in scripts/. Not a 4 because commands are complete with real arguments and cover local, branch, PR-base, commit, panel, and Windows/PowerShell variants.

5 / 5

Workflow Clarity

A clear operational sequence exists (set paths once → pick target → run → verify/rerun after fixes → final report) with explicit validation checkpoints and feedback loops ("rerun focused tests and rerun the structured review helper", "Stop as soon as the helper exits 0 with no accepted/actionable findings", two-cycle convergence pause). Not a 5 because the sequence is fragmented across many interleaved sections (Scope Governor, Oversized Bundles, Parallel Closeout, Models, Isolation) making the end-to-end flow harder to follow, and the ordering is not strictly linear.

4 / 5

Progressive Disclosure

Structure is present (clear ## sections, executable scripts correctly externalized to scripts/autoreview and scripts/test-review-harness, both real files), but large bodies of detail that belong in reference files are inlined in SKILL.md itself — the engine-isolation tables, model/thinking tables, environment-default tables, and per-engine edge-case policy. This matches 'some structure but content that should be separate is inline'; not a 4 because no references/ files exist and the one-level-deep reference pattern is largely unused, and not a 2 because section headers and the script indirection do keep the document navigable.

3 / 5

Total

14

/

20

Passed

Description

48%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is extremely terse: it identifies the domain and engine defaults but omits any trigger guidance and any enumeration of what the skill actually does. Distinctiveness is adequate, but completeness and specificity are held back by the missing 'Use when' clause and absent concrete actions.

Suggestions

Add an explicit trigger clause, e.g. 'Use when the user asks for a Codex/Claude/Pi review, a second-model review, or a review before committing or shipping changes.'

List 2-4 concrete actions so the 'what' goes beyond a domain label, e.g. building a validated review bundle, running the selected engine, verifying findings, and rerunning after fixes.

Include natural synonyms users would say ("second opinion", "review my changes", "pre-ship review") to improve trigger term coverage.

DimensionReasoningScore

Specificity

The description names the domain ("Pre-commit/ship code review") but lists no concrete actions or capabilities — it only identifies engines ("Codex default; optional Claude or Pi") rather than what the skill actually does. This matches the 'names the domain but actions are minimal or generic' anchor; it is not a 3 because there is no 1-2 item list of concrete actions, only a compressed noun phrase.

2 / 5

Completeness

The 'what' is stated with reasonable clarity (pre-commit/ship code review with a default engine), but there is no 'Use when...' clause or equivalent trigger guidance at all, which caps completeness at 3 per the rubric guideline. It fits the 'clear what but when is missing' anchor; not a 2 because the what is explicit and engine defaults are stated, and not a 4 because the when is entirely absent rather than merely imprecise.

3 / 5

Trigger Term Quality

It includes some relevant natural keywords ("Pre-commit", "ship", "code review", "Codex", "Claude", "Pi") but misses common variations a user would actually say ("review my changes", "review before commit/ship", "second-model review", "PR review"). This fits 'some relevant keywords but missing common variations or synonyms'; not a 2 because the present terms are domain-specific rather than generic, and not a 4 because coverage lacks natural trigger phrasing.

3 / 5

Distinctiveness Conflict Risk

The niche is mostly distinct — pre-commit/ship review tied to specific engines (Codex/Claude/Pi) — with minor overlap risk against generic code-review skills triggered by "code review" alone. Not a 5 because "code review" is a common trigger that could collide with other review skills, and not a 3 because the engine names and pre-commit/ship framing narrow it considerably.

4 / 5

Total

12

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
openclaw/clawhub
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.