CtrlK
BlogDocsLog inGet started
Tessl Logo

research-review

Get a deep critical review of research from GPT using a secondary Codex agent. Use when user says "review my research", "help me review", "get external review", or wants critical feedback on research ideas, papers, or experimental results.

58

Quality

67%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/skills-codex/research-review/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

60%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a genuinely actionable, well-sequenced workflow with strong prompt scaffolding for a multi-round external review. Its two real weaknesses are avoidable duplication that inflates token cost, and dangling shared-references that break the file's own progressive-disclosure architecture.

Suggestions

Create the referenced shared-references files (reviewer-routing.md, output-composition.md, review-tracing.md) or inline their essential rules and drop the links, so no reference dead-ends.

Delete or merge the 'Prompt Templates' section into the Step 2 prompt block and fold 'Key Rules' into the workflow steps to remove duplicated content.

Add a brief validation checkpoint after each round (confirm the agent response arrived, then save/document) and an explicit fallback recipe for the no-delegation case.

DimensionReasoningScore

Conciseness

Mostly efficient step-by-step prose, but there is measurable duplication and padding: the 'Prompt Templates' section restates the Step 2 prompt, 'Key Rules' repeats workflow points already stated, and a long inline 'SCOPE LIMITS' block mixes meta-policy with the review prompt. It fits 'Mostly efficient but includes some unnecessary explanation or could be tightened'; not 4 because the duplication is avoidable, not 2 because there is no filler explaining things Claude already knows.

3 / 5

Actionability

Concrete, near-copy-paste spawn_agent and send_input blocks with explicit model/effort parameters, plus literal follow-up phrasings ("If we reframe X as Y..."). It falls short of 5 only because key slots remain placeholders ("[Full research context + specific questions]", "[saved reviewer id from Step 2]") and the fallback path when delegation is disallowed is described only at a high level.

4 / 5

Workflow Clarity

Six clearly sequenced steps with a defined convergence criterion (Step 4) and per-round guidance make the process easy to follow, and the iterative dialogue is itself a feedback loop. It is not 5 because explicit validation checkpoints are absent — e.g., no step verifies the spawned agent succeeded, that responses were actually saved, or that the review document was written before updating memory.

4 / 5

Progressive Disclosure

Sections are well-labeled, but all three cross-file references (../shared-references/reviewer-routing.md, output-composition.md, review-tracing.md) do not exist anywhere in the bundle — no references/, scripts/, or assets/ directories are present — so navigation dead-ends. This matches 'Minimal structure; ... references are buried' territory: the routing/tracing/composition policy content those files should hold is either dangling or inlined in SKILL.md (the long SCOPE LIMITS block), rather than split out one level deep.

2 / 5

Total

13

/

20

Passed

Description

75%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A well-formed description with an explicit what/when structure and genuine quoted trigger phrases. Its weaknesses are modest: capability coverage is thin relative to the body, and one overly broad trigger creates overlap risk with general review skills.

Suggestions

Enumerate 2-3 more concrete capabilities (e.g., multi-round adversarial dialogue, mock NeurIPS/ICML reviews, minimal-experiment design) to raise specificity.

Qualify the generic trigger "help me review" (e.g., "help me review my research/paper") to reduce conflict with code- or document-review skills.

DimensionReasoningScore

Specificity

The description names the domain ("deep critical review of research") and one concrete action (obtaining it "from GPT using a secondary Codex agent"), but stops short of the fuller capability set the body provides (multi-round dialogue, mock reviews, experiment design). It matches the anchor 'Names domain and 1-2 concrete actions, but not comprehensive' — a 4 would require several listed specific actions.

3 / 5

Completeness

It explicitly answers what ("Get a deep critical review of research from GPT using a secondary Codex agent") and when ("Use when user says... or wants critical feedback on..."), with concrete quoted trigger phrases — the exact shape of the 5 anchor. Not a 4, because the 'when' clause is already explicit and specific rather than merely adequate.

5 / 5

Trigger Term Quality

Natural user phrases are quoted directly ("review my research", "help me review", "get external review") plus object coverage ("research ideas, papers, or experimental results"), which is good keyword coverage. It falls short of the 5 anchor because common variants like "critique my paper", "review my draft", or "is my work ready for submission" are missing.

4 / 5

Distinctiveness Conflict Risk

The research/paper-critique niche is fairly distinct, but the trigger "help me review" is generic and would plausibly fire on code-review or document-review requests, giving overlap risk with closely related skills. This fits 'Somewhat specific but could still overlap with similar skills'; a 4 would require the broad generic phrase to be qualified.

3 / 5

Total

15

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 1 suspicious

Warning

Total

15

/

16

Passed

Repository
wanshuiyin/Auto-claude-code-research-in-sleep
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.