CtrlK
BlogDocsLog inGet started
Tessl Logo

auto-review-loop-minimax

Autonomous multi-round research review loop using MiniMax API. Use when you want to use MiniMax instead of Codex MCP for external review. Trigger with "auto review loop minimax" or "minimax review".

60

Quality

71%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Critical

Do not install without reviewing

Fix and improve this skill with Tessl

tessl review fix ./skills/auto-review-loop-minimax/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

63%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-sequenced, highly actionable autonomous-loop protocol with strong state-recovery and documentation discipline. Its two real weaknesses are redundancy (the same API call and system prompt repeated three times) and inlining of prompt-template material that belongs in a reference file, plus no error-handling path for failed or unparseable reviewer responses.

Suggestions

Define the reviewer system prompt and the curl/MCP invocation once (e.g. in the API Configuration section) and reference it from Phase A and the Round 2+ template instead of repeating all three blocks verbatim — this alone would remove ~70 lines.

Move the Round 2+ prompt template to a references/ file (e.g. references/round-2-prompt.md) and keep only a one-line pointer plus the key fields (previous score, verdict, weaknesses, changes) in the body.

Add a checkpoint in Phase B for API failure: if curl fails, returns an HTTP error, or the response has no extractable score/verdict, retry once and otherwise record the failure in AUTO_REVIEW.md and continue to the next round rather than proceeding as if a review happened.

DimensionReasoningScore

Conciseness

Mostly efficient project-specific protocol that avoids explaining concepts Claude already knows, but the curl command with the identical system prompt ("You are a senior machine learning researcher...") is repeated three times — API Configuration, Phase A, and the Round 2+ template — and the MCP invocation block likewise appears twice. This ~70-line duplication is exactly what could be tightened into one referenced template.

3 / 5

Actionability

Concrete throughout: exact curl commands, the REVIEW_STATE.json schema, exact file paths and legacy fallbacks, an explicit stop condition ("score >= 6 AND verdict ∈ {ready, almost}"), and a fill-in log template. Minor gaps keep it from 5: prompt blocks remain placeholder templates rather than fully assembled examples, and Phase D ('Monitor remote sessions for completion') gives no concrete command.

4 / 5

Workflow Clarity

The Initialization → Phases A–E → Termination sequence is clear, with an explicit stop condition, resume/staleness rules (24h timestamp check), and prioritization rules — a genuine feedback loop. The missing checkpoints are around the API call itself: no handling for a failed curl, an HTTP error, or a response that doesn't contain a parseable score/verdict, which is a real fragility in a loop driven by parsing that response.

4 / 5

Progressive Disclosure

Sections are well organized and the ../shared-references/ links are clearly signaled, but no bundle files exist alongside this SKILL.md, so those four referenced protocols are unverifiable. Meanwhile ~55 lines of the 'Prompt Template for Round 2+' section — content that clearly belongs in a reference file — are inlined in the main body, matching the anchor of structure present but separable content inline.

3 / 5

Total

14

/

20

Passed

Description

78%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description that explicitly answers both what the skill does and when to use it, with concrete, natural trigger phrases and clear differentiation from the Codex MCP variant. The main gaps are the second-person voice and a 'what' clause that stays one level above the actual review→fix→re-review mechanics.

Suggestions

State the concrete loop mechanics in the 'what' clause, e.g. 'Autonomously runs review → implement fixes → re-review rounds against an external MiniMax reviewer until a positive assessment or max rounds.'

Rewrite the 'when' clause in third person, e.g. 'Use when the user wants an external MiniMax review instead of Codex MCP' — this removes the second-person phrasing penalized by the rubric.

Add one or two natural trigger variants users might actually say, such as 'review my research' or 'minimax reviewer'.

DimensionReasoningScore

Specificity

"Autonomous multi-round research review loop using MiniMax API" names the domain and 1-2 concrete actions (multi-round review, external review), but never states the actual loop mechanics (review → implement fixes → re-review) or the outputs produced. The second-person phrasing "Use when you want to use MiniMax" takes a further point off per the third-person-voice guideline.

3 / 5

Completeness

Both questions are answered explicitly: the 'what' ("Autonomous multi-round research review loop using MiniMax API") and the 'when' ("Use when you want to use MiniMax instead of Codex MCP for external review") plus concrete trigger phrases ("Trigger with 'auto review loop minimax' or 'minimax review'"). This matches the anchor-5 example's structure of capability statement + 'Use when' + explicit triggers.

5 / 5

Trigger Term Quality

Explicit trigger phrases "auto review loop minimax" and "minimax review" are natural phrasings for this niche, alongside "MiniMax" and "Codex MCP" as differentiating keywords. A few natural variations (e.g. "review my paper", "external review", "MiniMax-M3") are missing.

4 / 5

Distinctiveness Conflict Risk

The niche is clear and the MiniMax/Codex-MCP distinction plus dedicated trigger phrases keep it distinct, but it shares its core purpose (autonomous research review loop) with the closely related /auto-review-loop skill, leaving minor overlap risk on generic 'review loop' phrasings.

4 / 5

Total

16

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

relative_links

Relative link issues: 4 suspicious

Warning

Total

13

/

16

Passed

Repository
wanshuiyin/Auto-claude-code-research-in-sleep
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.