CtrlK
BlogDocsLog inGet started
Tessl Logo

auto-review-loop

Autonomous multi-round research review loop. Repeatedly reviews using a secondary Codex agent, implements fixes, and re-reviews until positive assessment or max rounds reached. Use when user says "auto review loop", "review until it passes", or wants autonomous iterative improvement.

68

Quality

83%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

77%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is exceptionally actionable and the workflow is superbly sequenced with validation checkpoints and recovery paths. Its weaknesses are token efficiency — heavy duplication of the scope-limits block and repeated rules — and monolithic inlining of material (test specs, protocols, templates) that belongs in reference files for progressive disclosure.

Suggestions

Replace the two verbatim 20-line SCOPE LIMITS blocks with a single pointer to ../shared-references/review-scope-limits.md (the file the skill already references for nightmare mode), saving ~40 lines.

Move the four Acquittal Gate test specifications and the Debate Protocol templates into a references/ file (e.g. references/acquittal-gate-tests.md), keeping only the operative rules in SKILL.md.

Consolidate the stop-condition and reviewer-memory rules, each currently stated in 3–4 places, into one authoritative statement with cross-references.

DimensionReasoningScore

Conciseness

The ~500-line body duplicates the entire 20-line SCOPE LIMITS block verbatim in both the medium-mode prompt and the round-2+ template (despite a shared-references file already existing for it), restates the stop condition in four places, and repeats memory/debate rules across the dedicated section and Phases B.5/B.6. Mostly efficient with no conceptual padding, but clearly could be tightened — anchor 3.

3 / 5

Actionability

Fully executable guidance throughout: exact spawn_agent/send_input blocks with model and reasoning effort pinned, complete JSON state schema, JSONL receipt format with field-level rules, run_id generation recipe, concrete render command with flags, and explicit prioritization rules for fixes. Copy-paste ready — anchor 5.

5 / 5

Workflow Clarity

The multi-phase workflow is clearly sequenced (Initialization decision tree covering fresh/resume/stale/legacy cases, Phases A–E, stop gate B.5.1 deliberately placed after the memory update), with explicit validation checkpoints, feedback loops (debate protocol, regression-triggered trace diffing), and four test specs validating the acquittal gate. Anchor 5.

5 / 5

Progressive Disclosure

No bundle files exist (references/, scripts/, assets/ are absent) and all referenced shared-references live outside the skill, unverifiable here. The body is well-sectioned but monolithic: the four Acquittal Gate test specifications, the full debate/memory protocols, and the duplicated prompt templates clearly belong in separate reference files — anchor 3 ('content that should be separate is inline').

3 / 5

Total

16

/

20

Passed

Description

88%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: concrete, third-person, and complete with explicit 'Use when...' triggers quoting natural user phrases. The only gaps are a few missing trigger synonyms and minor overlap risk with the family of related review skills.

Suggestions

Add one or two more natural trigger variants, e.g. 'keep reviewing until it passes' or 'iterate until the reviewer approves', to broaden trigger coverage.

Sharpen distinctiveness by naming the loop's distinguishing feature (autonomous review-fix-rereview cycle with an external reviewer) more prominently to reduce collision with single-pass review skills.

DimensionReasoningScore

Specificity

The description lists multiple concrete actions — 'Repeatedly reviews using a secondary Codex agent, implements fixes, and re-reviews until positive assessment or max rounds reached' — covering the loop's full behavior with no material gaps, matching the comprehensive-coverage anchor.

5 / 5

Completeness

It explicitly answers both 'what' (multi-round review → fix → re-review loop with a secondary Codex agent, bounded by positive assessment or max rounds) and 'when' with concrete quoted trigger phrases, exactly matching the anchor-5 example pattern.

5 / 5

Trigger Term Quality

It quotes verbatim natural phrases users would say ("auto review loop", "review until it passes") plus a paraphrase ("autonomous iterative improvement"), but misses common variants such as 'keep reviewing until it passes' or 'iterate until the reviewer approves'. This sits between good coverage (4) and comprehensive synonym coverage (5).

4 / 5

Distinctiveness Conflict Risk

The loop framing and named triggers ('auto review loop', 'review until it passes') form a clear niche, but the broad term 'reviews' and the crowded family of review/audit skills in this ecosystem create minor overlap risk with closely related skills — anchor 4 fits better than 5.

4 / 5

Total

18

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (501 lines); consider splitting into references/ and linking

Warning

relative_links

Relative link issues: 5 suspicious

Warning

Total

14

/

16

Passed

Repository
wanshuiyin/Auto-claude-code-research-in-sleep
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.