CtrlK
BlogDocsLog inGet started
Tessl Logo

auto-review-loop-llm

Autonomous research review loop using any OpenAI-compatible LLM API. Configure via llm-chat MCP server or environment variables. Trigger with "auto review loop llm" or "llm review".

53

Quality

61%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Critical

Do not install without reviewing

Fix and improve this skill with Tessl

tessl review fix ./skills/auto-review-loop-llm/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

63%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body presents a clear, well-sequenced loop (review → parse → fix → wait → document) with explicit stop conditions, state persistence for recovery, and mostly concrete MCP/curl guidance. Its main costs are duplicated curl and prompt-template blocks, changelog-style meta commentary, and provider/configuration bulk that belongs in a separate reference file — plus shared-reference links that cannot be verified against the bundle.

Suggestions

Deduplicate the curl fallback and reviewer prompt into one parameterized template, and delete the changelog-style note "(Earlier wording used `or` and a stale verdict set; the `AND` form is authoritative.)" in favor of the single authoritative statement.

Move the 9-provider configuration table and MCP settings example into a references/ file (e.g. references/providers.md) and link to it, keeping SKILL.md as a lean overview.

Add explicit error-handling checkpoints: what to do when the review response has no parseable score, when the MCP call or curl fails, and how many retries before aborting — this would also lift workflow clarity toward anchor 5.

DimensionReasoningScore

Conciseness

Quotes: the curl fallback appears three times ("## API Call Method", "#### Phase A" fallback, and again verbatim), the reviewer prompt template is repeated (Phase A and "## Prompt Template for Round 2+"), and changelog-style meta commentary is inlined ("(Earlier wording used `or` and a stale verdict set; the `AND` form is authoritative.)"). The content is mostly efficient and free of concept over-explanation, but the duplication and self-referential notes mean it could be tightened, matching anchor 3 ("mostly efficient but includes some unnecessary explanation or could be tightened").

3 / 5

Actionability

Quotes: concrete MCP invocations ("mcp__llm-chat__chat:" with prompt/model/system fields), a copy-paste curl command with the full JSON body ("curl -s "${LLM_BASE_URL}/chat/completions" ... -d '{"model": "${LLM_MODEL}", "messages": [...], "max_tokens": 4096}'"), an exact state-file schema, and a complete output markdown template. Placeholders like "[Full research context: claims, methods, results, known weaknesses]" are appropriate parameterization for a loop skill, leaving only minor gaps (e.g. no jq/python snippet for parsing the score out of the response), fitting anchor 4 ("mostly executable guidance; concrete code or commands with minor gaps").

4 / 5

Workflow Clarity

Quotes: "### Initialization ... 1. Check `review-stage/REVIEW_STATE.json` for recovery", "### Loop (up to MAX_ROUNDS)" with Phases A-E, an explicit STOP checkpoint ("**STOP**: If score >= 6 AND verdict ∈ {"ready", "almost"}"), and termination steps ("Set `review-stage/REVIEW_STATE.json` status to "completed""). The entire skill is a validate→fix→retry feedback loop with state recovery, matching anchor 4; it falls short of 5 only because there is no error-recovery checkpoint for API failures or unparseable review responses.

4 / 5

Progressive Disclosure

The body links four one-level-deep shared references ("[Output Versioning Protocol](../shared-references/output-versioning.md)", output-manifest, output-language, and external-cadence), but no references/, scripts/, or assets/ bundle files exist and the ../shared-references/*.md targets are not present in the bundle, so the links cannot be verified. Meanwhile substantial configuration bulk (a 9-provider table of base URLs and models, full curl bodies) is inlined in SKILL.md rather than split out, fitting anchor 3 ("some structure but could be better organized; ... content that should be separate is inline").

3 / 5

Total

14

/

20

Passed

Description

58%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is third-person, concise, and includes an explicit trigger clause, which puts it well above the bad examples. Its main weaknesses are that it describes the mechanism (LLM API, MCP config) rather than the actions the loop performs, offers only two narrow trigger phrases, and does not clearly differentiate itself from the closely related /auto-review-loop sibling skill.

Suggestions

State the loop's concrete actions in the description, e.g. "Scores research work 1-10 via an external LLM reviewer, implements the minimum fixes, and re-reviews until score >= 6/10 or MAX_ROUNDS", so the what is complete.

Broaden trigger terms with natural variations such as "review loop", "auto review", "iterative review", or "research feedback" instead of only the two exact phrases.

Add a differentiating clause (e.g. "generic LLM version of /auto-review-loop, works with any OpenAI-compatible provider") to reduce overlap risk with the sibling skill.

DimensionReasoningScore

Specificity

Quotes: "Autonomous research review loop using any OpenAI-compatible LLM API", "Configure via llm-chat MCP server or environment variables" — the domain and delivery mechanism are concrete, but the description never states the actions the loop performs (score work, implement fixes, re-review until threshold), matching anchor 3 ("names domain and 1-2 concrete actions, but not comprehensive") rather than anchor 4 which requires several specific actions listed.

3 / 5

Completeness

The "what" ("Autonomous research review loop using any OpenAI-compatible LLM API") is clear, and the "when" is explicit trigger guidance ("Trigger with...") so the completeness cap of 3 does not apply. It falls between anchor 4 ("'when' could be more explicit or specific") and anchor 5, because the trigger phrases are concrete but the "what" omits the loop's actual behavior and stop condition, so 4 is the best fit.

4 / 5

Trigger Term Quality

Quotes: "Trigger with 'auto review loop llm' or 'llm review'" — two explicit trigger phrases exist, but coverage is narrow and misses common variations users would naturally say ("review loop", "auto review", "iterative review", "research feedback"), fitting anchor 3 ("some relevant keywords but missing common variations or synonyms") rather than anchor 4's "good keyword coverage".

3 / 5

Distinctiveness Conflict Risk

Quotes: "Autonomous research review loop using any OpenAI-compatible LLM API" — the research-review-loop niche is somewhat specific, but the skill body itself references a closely related sibling skill ("Like /auto-review-loop") doing nearly the same job, and the generic trigger phrase "llm review" could plausibly fire for other LLM-related skills, matching anchor 3 ("somewhat specific but could still overlap with similar skills").

3 / 5

Total

13

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

relative_links

Relative link issues: 4 suspicious

Warning

Total

13

/

16

Passed

Repository
wanshuiyin/Auto-claude-code-research-in-sleep
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.