CtrlK
BlogDocsLog inGet started
Tessl Logo

research-review

Get a deep critical review of research from an external reviewer backend (Codex or manual). Use when user says "review my research", "help me review", "get external review", or wants critical feedback on research ideas, papers, or experimental results.

63

Quality

76%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/research-review/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

73%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An instruction-only orchestration skill with excellent workflow structure — sequenced steps, explicit convergence criteria, and real error paths — and very concrete MCP calling conventions. Its weaknesses are repetition (manual config stated twice, Key Rules restating constants), an off-topic policy block inlined into the reviewer prompt, and a missing review-brief template for the skill's central artifact.

Suggestions

State the manual-review config dict once (in the Reviewer Calling Convention) and reference it from Step 2/Key Rules instead of repeating it verbatim, and drop Key Rules bullets that restate the Constants section.

Move the ~20-line SCOPE LIMITS policy block out of the inline reviewer prompt into a shared-reference file (it is policy for the reviewer, not procedure for the executor), keeping only a one-line pointer in the prompt.

Add a short review-brief template (context, core claims, methodology, key results, known weaknesses, specific questions, artifact paths) so the skill's central artifact is as copy-paste ready as the MCP calls.

DimensionReasoningScore

Conciseness

Mostly purposeful, but real tightening is needed: the manual-backend config dict appears verbatim twice (Reviewer Calling Convention lines 34-41), Key Rules restates model pinning already given in Constants and Step 2, and the ~20-line 'SCOPE LIMITS' policy block inside the reviewer prompt is off-topic for the skill's core job. This exceeds anchor 4's 'minor instances of over-explanation', fitting anchor 3's 'some unnecessary explanation or could be tightened'.

3 / 5

Actionability

Highly concrete: exact MCP tool names, exact model/config values (gpt-6-astra, {"model_reasoning_effort": "ultra"}), a copy-paste-ready reviewer prompt, threadId reuse rules, and specific follow-up phrasings. Not 5 because the review brief — the central artifact — has no template or structure (one descriptive sentence only), and full execution (tracing, routing, composition) depends on shared-references files and save_trace.sh that are not present in the bundle.

4 / 5

Workflow Clarity

Clear 5-step sequence (gather context → initial review → iterative dialogue → convergence → document) with an explicit convergence checklist, feedback rounds that check 'whether the revision actually fixed them', and explicit error paths: stop and print the install command if manual-review MCP is unavailable, and emit REVIEW_UNAVAILABLE rather than guessing on a missing reviewer identity. Matches anchor 5's explicit validation and feedback loops.

5 / 5

Progressive Disclosure

Well-organized sections with clearly signaled, one-level-deep references to shared-references files (bold names with inline links). Not 5 because content that belongs in those references is inlined — the ~20-line SCOPE LIMITS policy block and ~10 lines of composed-mode rules despite 'Full rules: output-composition.md' — and the referenced ../shared-references/*.md files and save_trace.sh are not present in the bundle to verify. Clearly above anchor 3's 'references not clearly signaled'.

4 / 5

Total

16

/

20

Passed

Description

78%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description with explicit what/when structure and natural trigger phrases in correct third-person voice. Its main weakness is underselling the skill's concrete capabilities — multi-round reviewer dialogue, experiment design, claims matrices, and mock venue reviews — which would also sharpen distinctiveness against generic review skills.

Suggestions

Enumerate the skill's concrete capabilities in the description (e.g., 'Runs multi-round adversarial review dialogue, designs minimal experiment packages, builds results-to-claims matrices, and writes mock NeurIPS/ICML reviews') to lift specificity from one action to several.

Add distinguishing trigger phrases that separate this from generic review skills, e.g., 'external review', 'cross-model review', 'mock reviewer score', so it does not fire for plain code or document review requests.

DimensionReasoningScore

Specificity

Names the domain and one concrete action ("Get a deep critical review of research from an external reviewer backend (Codex or manual)") with a concrete mechanism, but does not enumerate the skill's several capabilities (multi-round dialogue, experiment design, claims matrix, mock reviews). Not 4 because 'several specific actions' are absent; not 2 because the stated action and medium are concrete rather than generic.

3 / 5

Completeness

Explicitly answers both what ("Get a deep critical review of research from an external reviewer backend") and when ("Use when user says 'review my research', 'help me review', 'get external review', or wants critical feedback on research ideas, papers, or experimental results") with concrete quoted trigger phrases, in third-person voice. Matches the anchor-5 example pattern exactly.

5 / 5

Trigger Term Quality

Includes natural phrases users would say ("review my research", "help me review", "get external review") plus domain objects ("research ideas, papers, or experimental results"). Not 5 because common variants like "critique my paper/draft" or "mock review" are missing; clearly above 3's partial keyword coverage.

4 / 5

Distinctiveness Conflict Risk

The external-reviewer-backend framing carves a clear niche, but "help me review" and "review my research" could also plausibly match generic review or code-review skills. Minor overlap risk with closely related skills fits anchor 4; not 5 given that overlap, not 3 since the niche is genuinely distinct.

4 / 5

Total

16

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

relative_links

Relative link issues: 2 suspicious

Warning

Total

13

/

16

Passed

Repository
wanshuiyin/Auto-claude-code-research-in-sleep
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.