CtrlK
BlogDocsLog inGet started
Tessl Logo

research-review

Get a deep critical review of research from Gemini via gemini-review MCP. Use when user says "review my research", "help me review", "get external review", or wants critical feedback on research ideas, papers, or experimental results.

60

Quality

71%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/skills-codex-gemini-review/research-review/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

63%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The content delivers a genuinely actionable multi-round review workflow with concrete tool names, commands, and polling mechanics. It is dragged down by an off-topic ~25-line scope-limits block inside the Round 1 prompt, duplicated prompt templates, and no progressive disclosure — all templates and prompt text live inline in one long file with no reference bundle.

Suggestions

Move the five Prompt Templates into a single reference file (e.g. references/prompt-templates.md) and keep one representative template inline in Step 2, cutting the SKILL.md body substantially.

Replace the ~25-line SCOPE LIMITS block in the Round 1 prompt with one or two sentences of scope guidance actually relevant to research review (e.g. 'focus on logical gaps and missing experiments; do not manufacture findings'), or move it to a reference file — it currently pads the highest-token-cost part of the skill.

Add an explicit error-recovery checkpoint to the workflow: what to do when review_status never reaches done=true, when a round's response is empty, or when threadId resumption fails, mirroring the existing polling loop.

DimensionReasoningScore

Conciseness

The workflow steps are written as lean imperatives, but the ~25-line "SCOPE LIMITS" block embedded in the Round 1 prompt discusses threat models, hash/fingerprint schemes, and feature-flag aversion — content unrelated to research review — and the trailing Prompt Templates section restates prompts already given inline. This is 'mostly efficient but includes some unnecessary explanation or could be tightened'; not a 2 because the bulk of the body is functional instruction rather than padding.

3 / 5

Actionability

Concrete, executable guidance dominates: exact MCP tool names (mcp__gemini-review__review_start, review_reply_start, review_status), a runnable install command (`codex mcp add gemini-review -- python3 ~/.codex/mcp-servers/gemini-review/server.py`), polling instructions with waitSeconds/done=true, and copy-ready follow-up phrasings. Minor gaps keep it below 5: the exact call syntax for reply_start is never shown and the Round 1 prompt is a placeholder template ('[Full research context + specific questions]').

4 / 5

Workflow Clarity

A clearly sequenced 5-step workflow (gather context → round 1 → iterative rounds → convergence → document) with explicit state management (save jobId, poll until done=true, retain threadId) and defined convergence criteria. It sits at anchor 4 rather than 5 because there is no error-recovery checkpoint for reviewer failures/timeouts, and no validation that the saved review document is complete before finishing.

4 / 5

Progressive Disclosure

The body has clear section headers but is a single ~135-line file with no bundle files (references/, scripts/, assets/ do not exist), so everything — the large scope-limits prompt text, the follow-up patterns, and the five prompt templates — is inlined in SKILL.md. This matches 'Some structure but could be better organized; content that should be separate is inline' rather than 4, since the prompt-template library in particular belongs in a one-level-deep reference file.

3 / 5

Total

14

/

20

Passed

Description

78%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description that explicitly answers both what the skill does and when to use it, with quoted natural trigger phrases and a distinctive mechanism (Gemini via gemini-review MCP). Its main weakness is that the 'what' is a single action, leaving the multi-round review, mock-review, and experiment-design capabilities undescribed.

DimensionReasoningScore

Specificity

The description names the domain and one concrete action — "Get a deep critical review of research from Gemini via gemini-review MCP" — but does not enumerate the skill's other capabilities (multi-round dialogue, mock reviews, experiment design). This matches the anchor 'Names domain and 1-2 concrete actions, but not comprehensive'; a 4 would require several specific actions listed.

3 / 5

Completeness

Both halves are explicit: what — "Get a deep critical review of research from Gemini via gemini-review MCP"; when — "Use when user says 'review my research', 'help me review', 'get external review', or wants critical feedback on research ideas, papers, or experimental results". This mirrors the anchor-5 example structure (clear what plus concrete quoted trigger phrases), so it is not the 4 anchor where the 'when' is only loosely specified.

5 / 5

Trigger Term Quality

Natural trigger phrases are quoted directly ("review my research", "help me review", "get external review") plus "critical feedback on research ideas, papers, or experimental results". Coverage is good but misses common variations like 'review my paper', 'is this publishable', or 'mock review'. Fits 'Good keyword coverage; a few natural terms missing' rather than the comprehensive synonym coverage of a 5.

4 / 5

Distinctiveness Conflict Risk

Naming Gemini and the gemini-review MCP carves a clear niche (external cross-family review of research), distinct from generic review skills. Minor overlap remains: "help me review" and "critical feedback" could plausibly collide with code-review or document-review skills. Fits 'Mostly distinct; minor overlap risk' — not a 5 because the generic 'help me review' trigger is broad.

4 / 5

Total

16

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
wanshuiyin/Auto-claude-code-research-in-sleep
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.