CtrlK
BlogDocsLog inGet started
Tessl Logo

auto-review-loop-llm

Autonomous research review loop using any OpenAI-compatible LLM API. Configure via llm-chat MCP server or environment variables. Trigger with "auto review loop llm" or "llm review".

55

Quality

63%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./skills/skills-codex/auto-review-loop-llm/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

63%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body delivers a genuinely executable autonomous loop: concrete MCP/curl calls, explicit stop conditions, state persistence, and documentation templates. Its weaknesses are duplicated prompt/curl blocks and stale editorial notes that pad the token budget, inlined configuration/provider content that belongs in reference files, and missing error-recovery checkpoints for API or parse failures.

Suggestions

Deduplicate the review prompt and curl fallback: define each once (e.g., in a single Prompt Template section) and reference it from Phase A, and delete the stale-versioning parenthetical in the POSITIVE_THRESHOLD constant, which is editing history rather than instruction.

Move the provider table and MCP server setup JSON into a reference file (e.g., references/llm-providers.md) and keep SKILL.md to the loop workflow, or trim the provider list to the one or two providers actually used.

Add explicit error-recovery checkpoints in Phase B for API failures and unparseable responses (e.g., retry via curl fallback, or re-prompt for a numeric score) so the workflow's validation loop covers failure modes, not just the happy path.

DimensionReasoningScore

Conciseness

The review prompt appears near-verbatim in both Phase A and the "Prompt Template for Round 2+" section, the curl fallback is duplicated, and the parenthetical "(Earlier wording used \"or\" + a stale verdict set; the AND form is authoritative.)" is revision history rather than instruction — mostly efficient but with unnecessary duplication that should be tightened, matching the 3 anchor rather than the 4 anchor's "minor instances".

3 / 5

Actionability

Concrete executable material is provided: the mcp__llm-chat__chat call syntax, complete curl commands with JSON bodies, an exact REVIEW_STATE.json schema, and a copy-paste markdown template for AUTO_REVIEW.md. Minor gaps remain — "Priority: metric additions > reframing > new experiments" and "Monitor remote experiments" are one-line directions with no commands — so it fits the mostly-executable 4 anchor rather than the fully copy-paste-ready 5.

4 / 5

Workflow Clarity

The sequence (Initialization → Phases A–E → Termination) is explicit with a precise STOP condition ("If score >= 6 AND verdict ∈ {\"ready\", \"almost\"}") and state persistence for crash recovery, and the loop itself is a validate→fix→retry feedback loop. It falls short of 5 because there are no checkpoints for API or parse failures (e.g., what to do if curl errors or the score cannot be extracted).

4 / 5

Progressive Disclosure

No bundle files (references/, scripts/, assets/) exist, and the body inlines content that would fit a reference file (the 8-row provider table, MCP server setup config) in a ~240-line SKILL.md. The only external links point outside the skill bundle to ../../shared-references/*.md, which cannot be verified from the skill directory, so structure is present but organization and signaling leave clear room for improvement.

3 / 5

Total

14

/

20

Passed

Description

62%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description clearly names its niche (autonomous research review loop over an OpenAI-compatible API) and includes explicit trigger phrases, so both what and when are answered. It is held back by thin action specificity in the description itself and narrow trigger-term coverage that misses natural user phrasings.

Suggestions

List the concrete loop actions in the description itself (e.g., "Runs review → implement fixes → re-review rounds with an external LLM reviewer until the work is submission-ready") to raise specificity.

Broaden trigger terms to natural user phrasings such as "use when the user wants an LLM reviewer to iteratively critique and improve research work" in addition to the exact trigger strings.

Rewrite the when-clause as a situation-based "Use when..." statement rather than only quoting the trigger keywords, which would make the trigger guidance more explicit.

DimensionReasoningScore

Specificity

"Autonomous research review loop" and "Configure via llm-chat MCP server or environment variables" name the domain and configuration mechanism, but the description lists only the single loop action rather than several concrete actions (review, implement fixes, re-review), matching the anchor for domain plus 1-2 concrete actions without comprehensive coverage.

3 / 5

Completeness

The description answers both what ("Autonomous research review loop using any OpenAI-compatible LLM API") and when ("Trigger with \"auto review loop llm\" or \"llm review\""), but the when-clause is meta-instructional trigger phrasing rather than a situation-based "Use when..." clause, so it fits the anchor for both present with room for more explicitness rather than the clearly-explicit 5 anchor.

4 / 5

Trigger Term Quality

The explicit triggers "auto review loop llm" and "llm review" give two relevant keywords, but common natural variations users would say ("review my paper", "reviewer feedback", "LLM reviewer") are missing, matching the anchor for some relevant keywords but missing synonyms.

3 / 5

Distinctiveness Conflict Risk

"research review loop", "llm-chat MCP server", and "OpenAI-compatible LLM API" define a clear niche, though the short generic trigger "llm review" carries minor collision risk with general LLM-review or code-review skills, fitting the mostly-distinct anchor rather than the minimal-conflict 5.

4 / 5

Total

14

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 3 suspicious

Warning

Total

15

/

16

Passed

Repository
wanshuiyin/Auto-claude-code-research-in-sleep
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.