CtrlK
BlogDocsLog inGet started
Tessl Logo

auto-review-loop

Autonomous multi-round research review loop. Repeatedly reviews using Gemini via gemini-review MCP, implements fixes, and re-reviews until positive assessment or max rounds reached. Use when user says "auto review loop", "review until it passes", or wants autonomous iterative improvement.

66

Quality

80%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Critical

Do not install without reviewing

SKILL.md
Quality
Evals
Security

Quality

Content

77%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a highly actionable, well-sequenced autonomous workflow with real validation checkpoints and state recovery — workflow clarity and actionability are exemplary. Its weaknesses are token efficiency (verbatim duplicated SCOPE LIMITS blocks and polling instructions, editorial commentary in constants) and progressive disclosure (a long monolithic file whose prompt templates and scope rules should live in reference files).

Suggestions

Factor the ~20-line SCOPE LIMITS block, currently duplicated verbatim in the Phase A and Round 2+ prompt templates, into a single named constant or a referenced file to cut ~20 lines of pure duplication.

Move the full Phase A and Round 2+ prompt templates into a references/ file (e.g. references/prompts.md), keeping only the tool names and polling contract inline in SKILL.md.

Delete the editorial meta-commentary from the POSITIVE_THRESHOLD constant ("Earlier wording used 'or' + a stale verdict set...") — it documents revision history, not the rule itself.

DimensionReasoningScore

Conciseness

Mostly instruction-dense rather than explanatory, but it carries real padding: the ~20-line SCOPE LIMITS block is duplicated verbatim in both the Phase A and Round 2+ templates, the MCP polling instructions are repeated word-for-word, and the POSITIVE_THRESHOLD constant embeds editorial meta-commentary ("Earlier wording used 'or' + a stale verdict set; the AND form is authoritative"). It is not 2 because nearly all content is genuine operational guidance rather than concepts Claude already knows, and not 4 because the duplicated blocks are substantial removable tokens.

3 / 5

Actionability

Fully executable throughout: exact MCP tool names (mcp__gemini-review__review_start / review_status / review_reply_start) with jobId/threadId handling, a concrete JSON schema for REVIEW_STATE.json, a copy-paste markdown template for documenting rounds, a literal human-checkpoint message template, and explicit input-parsing rules ("go"/"continue"/"skip 1,3"/"stop"). This matches the 5 anchor — copy-paste-ready commands and templates covering the common cases.

5 / 5

Workflow Clarity

The sequence is explicit — Initialization (with resume/stale-state validation branches) → Loop Phases A-E → Termination — with an exact STOP CONDITION (score >= 6 AND verdict in {"ready","almost"}), a state file written every Phase E for recovery, and a true feedback loop (review → implement fixes → re-review). This matches the 5 anchor: clear sequencing with explicit validation checkpoints and error-recovery handling.

5 / 5

Progressive Disclosure

The body is well-sectioned with clear headers, but it is a single ~300-line file with no bundle files (references/, scripts/, assets/ absent), and large blocks that belong in a reference are inlined — the duplicated SCOPE LIMITS prompt text and both full prompt templates. The Output Protocols links to ../../shared-references/*.md are one level deep but those files are not present in this bundle, so navigation cannot be verified. It is not 2 because structure and signaling are decent, and not 4 because significant content that should be split out is inline.

3 / 5

Total

16

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: it states what the skill does concretely, gives an explicit 'Use when...' clause with natural quoted trigger phrases, and anchors a distinct niche via the Gemini reviewer MCP. The main gap is modest trigger-synonym coverage and unmentioned capabilities (state persistence, notifications) that keep specificity and trigger quality just below top marks.

DimensionReasoningScore

Specificity

The description lists several concrete actions — "Repeatedly reviews using Gemini via gemini-review MCP, implements fixes, and re-reviews until positive assessment or max rounds reached" — which is concrete and actionable. It is not 5 because coverage is partial (state persistence, checkpoints, and notifications are unmentioned), and not 3 because it goes well beyond 1-2 generic actions to name the mechanism, the reviewer, and the termination condition.

4 / 5

Completeness

It explicitly answers both questions: what ("Autonomous multi-round research review loop... implements fixes, and re-reviews until positive assessment or max rounds reached") and when ("Use when user says 'auto review loop', 'review until it passes', or wants autonomous iterative improvement") with concrete trigger phrases. This matches the 5 anchor exactly; neither what nor when is vague or implied.

5 / 5

Trigger Term Quality

Triggers are natural phrases a user would actually say, quoted verbatim: "auto review loop", "review until it passes", "wants autonomous iterative improvement". It is not 5 because common synonyms and variations (e.g. "iterate until it passes", "keep reviewing", "make the review pass") are missing, and not 3 because multiple genuine user-facing phrases are present rather than generic keywords only.

4 / 5

Distinctiveness Conflict Risk

Naming "Gemini via gemini-review MCP" and "auto review loop" carves a clear niche that is unlikely to fire for unrelated skills, though the broad words "review" and "iterative improvement" leave minor overlap risk with general code-review/PR-review skills. It is not 5 because of that overlap, and not 3 because the MCP bridge and quoted trigger phrases make it substantially distinguishable.

4 / 5

Total

17

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 3 suspicious

Warning

Total

15

/

16

Passed

Repository
wanshuiyin/Auto-claude-code-research-in-sleep
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.