CtrlK
BlogDocsLog inGet started
Tessl Logo

auto-review-loop

Autonomous multi-round research review loop. In Copilot CLI it defaults to the native complementary rubber-duck subagent with host-event model evidence; elsewhere it uses Codex, while explicit external reviewer overrides remain available. Implements fixes and re-reviews until a policy-approved positive assessment or max rounds is reached.

51

Quality

58%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/auto-review-loop/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

63%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

This is an unusually actionable skill: every phase has executable commands, exact artifact schemas, and fail-closed validation. Its weaknesses are bloat and structure — massive verbatim duplication and routing rules inlined that belong in the referenced shared files — which both burn context and blur the control flow.

Suggestions

Define the SCOPE LIMITS block once (e.g. in a shared reference or a single section) and reference it from the four prompt templates instead of repeating it verbatim four times.

Move the full backend routing tables and copilot contract out of the body into shared-references/reviewer-routing.md, keeping only the routing summary in SKILL.md — the rules are currently stated in both places and can drift apart.

Delete changelog-style commentary (the "earlier wording used `or`" note) or relocate it to a deprecated/changes section so the authoritative rule reads cleanly.

DimensionReasoningScore

Conciseness

The ~1100-line body is heavily padded by duplication: the identical 23-line "SCOPE LIMITS" block is repeated verbatim four times, backend routing rules are restated across Constants, the Reviewer Calling Convention, Phase A, and Phase B.5.1, and changelog-style notes ("Earlier wording here used `or` and a stale verdict set ... that was an internal inconsistency") are inline. It does not explain concepts Claude already knows, so it stays above the lowest anchor.

2 / 5

Actionability

Guidance is fully executable: copy-paste-ready bash for the stop gate (with fail-closed REVIEW_UNAVAILABLE exits and helper resolution), the mktemp/heredoc rebuttal prompt construction, complete codex exec commands, exact MCP call shapes with config JSON, and exact file paths, artifact naming patterns, and JSON schemas for every state file.

5 / 5

Workflow Clarity

The loop is explicitly sequenced (Step -1 → Step 0 → Phases A–E → Termination) with a hard validation checkpoint (review_gate.py fail-closed gate, evidence revalidation) and an explicit stop condition plus a test-case checklist. It falls short of a 5 because the round_backend vs forward-looking REVIEWER_BACKEND distinction and escalation snapshot rules are explained in multiple places, making the actual control flow genuinely hard to trace without careful re-reading.

4 / 5

Progressive Disclosure

References are clearly signaled and one level deep (shared-references/reviewer-routing.md, external-cadence.md, review-tracing.md, tools/review_gate.py), but no bundle files ship with the skill while large blocks that clearly belong in those references are inlined anyway — four full reviewer prompt templates and the full backend routing tables that reviewer-routing.md supposedly already contains.

3 / 5

Total

14

/

20

Passed

Description

53%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description states the loop mechanics and stop condition concretely, but is written in internal-ecosystem jargon rather than user-facing language and entirely omits any 'Use when' trigger guidance. It is distinct from generic skills yet would not be reliably surfaced by a natural user request.

Suggestions

Add an explicit trigger clause, e.g. "Use when the user asks to review, critique, or iteratively improve research work, experiments, or a paper draft until an independent reviewer approves it."

Rewrite the jargon-heavy middle sentence in capability terms (e.g. "runs an independent cross-model reviewer of the work, then implements its fixes") instead of "native complementary rubber-duck subagent with host-event model evidence".

Include natural synonyms users would say ("review loop", "critique my results", "improve my paper") to improve trigger-term coverage.

DimensionReasoningScore

Specificity

Concrete actions are stated — "Implements fixes and re-reviews until a policy-approved positive assessment or max rounds is reached" — and the domain is named ("Autonomous multi-round research review loop"), but coverage is not comprehensive and the middle sentence is implementation jargon ("native complementary rubber-duck subagent with host-event model evidence") rather than user-facing capability.

3 / 5

Completeness

The 'what' is reasonably clear (autonomous review → fix → re-review loop with an explicit stop condition), but there is no "Use when..." clause or any equivalent trigger guidance, which caps completeness at 3 per the rubric.

3 / 5

Trigger Term Quality

Some relevant natural keywords appear ("research review", "review loop", "implements fixes"), but common variations and synonyms a user would actually say ("review my paper", "improve my research", "critique my experiments") are missing.

3 / 5

Distinctiveness Conflict Risk

The niche is fairly distinct — a multi-round adversarial cross-model research review loop with named backends (Codex, rubber-duck) — so conflict risk is limited to closely related review/critique skills, e.g. a one-shot review skill could plausibly overlap.

4 / 5

Total

13

/

20

Passed

Validation

75%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 12 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (1138 lines); consider splitting into references/ and linking

Warning

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

relative_links

Relative link issues: 5 suspicious

Warning

Total

12

/

16

Passed

Repository
wanshuiyin/Auto-claude-code-research-in-sleep
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.