CtrlK
BlogDocsLog inGet started
Tessl Logo

research-review

Get a deep critical review of research from Claude via claude-review MCP. Use when user says "review my research", "help me review", "get external review", or wants critical feedback on research ideas, papers, or experimental results.

64

Quality

76%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/skills-codex-claude-review/research-review/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

67%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The content is highly actionable — exact MCP tool names, a copy-paste install command, and complete prompt templates — organized into a clear six-step workflow with polling-based validation and explicit convergence criteria. Its weaknesses are token efficiency (duplicated assurance and polling text, a long embedded scope-limits policy block) and missing error-recovery guidance for stalled or failed review jobs.

Suggestions

Remove the duplicated cross-family assurance (the second blockquote after the title restates the opening block) and replace the verbatim-repeated polling instruction in Step 3 with a reference to Step 2's procedure.

Move the ~25-line SCOPE LIMITS policy block and the Prompt Templates section into a reference file (e.g., references/reviewer-prompts.md) and link to it, cutting SKILL.md's token weight while keeping the templates copy-pasteable.

Add a concrete review_status polling example with a suggested bounded waitSeconds value and a failure path (what to do if a job errors or never reaches done=true), which would close the workflow's main validation gap.

DimensionReasoningScore

Conciseness

The body is mostly operational, but there is noticeable padding that could be tightened: the cross-family assurance appears twice (the opening blockquote and "this route is a different model family from the Codex executor and records review_independence: cross-family" restated in the overlay assurance note), the polling instruction "immediately save the returned jobId and poll mcp__claude-review__review_status with a bounded waitSeconds until done=true" is repeated verbatim in Steps 2 and 3, and the ~25-line "SCOPE LIMITS" policy block plus the Prompt Templates section restate material. Anchor 3 (mostly efficient but could be tightened) fits; not 2 because there is no explanation of concepts Claude already knows and most tokens are functional.

3 / 5

Actionability

Concrete, executable guidance throughout: the install command `codex mcp add claude-review -- python3 ~/.codex/mcp-servers/claude-review/server.py`, exact tool names (`mcp__claude-review__review_start`, `review_reply_start`, `review_status`), fully written prompt bodies, and `tools: "Read,Grep,Glob"`. Minor gaps keep it below anchor 5: `threadId: [saved reviewer id from Step 2]` and `[Full research context + specific questions]` are placeholders, and there is no concrete example of the review_status polling call or a suggested waitSeconds value. Well above anchor 3 — this is real copy-paste-ready material, not pseudocode.

4 / 5

Workflow Clarity

Six clearly sequenced steps (gather context → initial review → iterative dialogue → convergence → document → trace) with explicit validation checkpoints ("poll mcp__claude-review__review_status with a bounded waitSeconds until done=true") and explicit stop criteria in Step 4. Anchor 4 fits: minor validation gaps — no guidance on what to do if a job errors, stalls, or never reaches done=true. Not 5 because there is no error-recovery feedback loop; not 3 because polling validation and convergence conditions are explicit, not implicit.

4 / 5

Progressive Disclosure

The body is well-sectioned (Constants, Prerequisites, Workflow, Key Rules, Prompt Templates) and the two references ([output-composition.md](../shared-references/output-composition.md) and ../shared-references/review-tracing.md) are clearly signaled one level deep. Anchor 4 fits: minor organization gaps — the referenced shared files live outside this skill's directory with no bundle files shipped alongside, and the long SCOPE LIMITS block and prompt templates are inlined where a reference file would reduce SKILL.md weight. Not 3 because the references that do exist are clearly signaled, and the overall structure is good; not 5 because content that could be split out remains inline and no in-bundle reference organization exists.

4 / 5

Total

15

/

20

Passed

Description

86%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description with an explicit what/when structure and natural user-voice trigger phrases. Its main limitation is breadth of stated actions — it advertises only the initial critical review, not the iterative follow-up rounds and deliverables (experiment designs, claims matrices, mock reviews) the skill provides.

DimensionReasoningScore

Specificity

"Get a deep critical review of research" names the domain and one concrete action applied to targets ("research ideas, papers, or experimental results"), but only that single review action is stated — not the multi-round dialogue, experiment design, or mock-review capabilities the skill actually has. This matches anchor 3 (domain and 1-2 concrete actions, not comprehensive), not anchor 4 which requires several listed specific actions.

3 / 5

Completeness

Explicitly answers both: what ("Get a deep critical review of research from Claude via claude-review MCP") and when ("Use when user says \"review my research\"... or wants critical feedback on research ideas, papers, or experimental results") with concrete trigger phrases. Matches anchor 5 exactly; not 4 because the when-clause is fully explicit, not merely present.

5 / 5

Trigger Term Quality

Direct user-voice phrases "review my research", "help me review", "get external review", plus synonyms "critical feedback" and target terms "research ideas, papers, or experimental results" give comprehensive natural-term coverage. Not 4: it captures the actual sentences a user would say, including variations, rather than leaving common synonyms missing.

5 / 5

Distinctiveness Conflict Risk

The research-review niche and claude-review MCP routing are distinct, but "help me review" and "get external review" could also match generic code-review or PR-review skills. Anchor 4 (mostly distinct, minor overlap risk with closely related skills) fits; not 5 because of that overlap with general review skills, and not 3 since the research framing is consistent throughout.

4 / 5

Total

17

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 1 suspicious

Warning

Total

15

/

16

Passed

Repository
wanshuiyin/Auto-claude-code-research-in-sleep
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.