CtrlK
BlogDocsLog inGet started
Tessl Logo

agentsociety-paper-review

Use when a completed research manuscript needs a robust internal, venue-calibrated mock peer review. Orchestrates three isolated reviewer subagents in parallel, then a separate MetaReview subagent that verifies evidence, reports agreement and score dispersion, and emits advisory reroutes. Never revises research artifacts, executes experiments, or updates pipeline state.

67

Quality

81%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

85%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-engineered orchestration skill: the workflow is fully sequenced with explicit validation gates and retry rules, references are real and one level deep, and the guidance is concrete throughout. The only deductions are minor — some deliberate redundancy between Workflow and Hard Constraints, and a missing example invocation for the PDF preparation script.

DimensionReasoningScore

Conciseness

The body is dense and assumes Claude's competence — no space is spent explaining peer review, PDFs, or hashing concepts, and sections like Venue Selection and Artifact Intake are pure operational guidance. It falls short of 5 because critical rules (digest recomputation, PDF intake gating, no cross-reviewer exposure) are stated twice, once in Workflow steps and again nearly verbatim in Hard Constraints, which could be tightened to single pointers.

4 / 5

Actionability

Guidance is highly concrete for an instruction-only skill: exact paths ('paper/reviews/<venue>-review-r<N>/'), a closed reroute vocabulary with eight literal values, a venue normalization table, manifest fields, and a real script ('scripts/prepare_pdf_review.py', verified to exist). It is not a 5 because the PDF step never gives the script's actual invocation (arguments, output flags), forcing the reader to open the script to run it — a minor but real executable gap.

4 / 5

Workflow Clarity

The eight-step Workflow is an unambiguous sequence with explicit validation checkpoints and feedback loops: 'Stop before dispatch when its status is failed', 'Retry a failed reviewer at most once under the protocol; never silently reduce the ensemble size', digest recomputation before MetaReview with round invalidation on mismatch, and dispatch gating on 'pass' or 'pass_with_warnings'. This matches the 5 anchor's clear sequence, explicit validation, and error-recovery loops; nothing is missing.

5 / 5

Progressive Disclosure

The SKILL.md body is a genuine overview with a 'Required References' section that clearly signals one-level-deep files, all of which exist in the bundle (ensemble-protocol.md, review-contract.md, meta-review-contract.md, reroute-guide.md, pdf-intake-contract.md, and six venue templates under references/venues/). The venues/ subdirectory is a flat grouping, not a reference chain, and detailed contracts are appropriately split out rather than inlined — matching the 5 anchor's well-signaled, easy navigation.

5 / 5

Total

18

/

20

Passed

Description

78%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, third-person description with an explicit 'Use when' clause, concrete multi-agent actions, and effective boundary disclaimers. Its main weakness is keyword breadth: it consistently uses 'manuscript' where most users would say 'paper' or 'paper draft', and it omits the venue names its own templates support.

Suggestions

Add 'paper' or 'paper draft' as a synonym for 'research manuscript' in the trigger clause (e.g., 'Use when a completed paper draft or research manuscript needs...') so the most natural user phrasing matches.

Name one or two supported venues (e.g., 'ICLR, NeurIPS, or another venue') in the description to catch users who ask for a venue-calibrated review by conference name.

Mention the concrete output (an ensemble review bundle with agreement analysis and reroutes) so the 'what' side is fully concrete.

DimensionReasoningScore

Specificity

The description lists several concrete actions — 'Orchestrates three isolated reviewer subagents in parallel, then a separate MetaReview subagent that verifies evidence, reports agreement and score dispersion, and emits advisory reroutes' — all in third person. It falls just short of the 5 anchor because the actions are orchestration-level and never name the concrete output artifact (e.g., an ensemble review bundle), leaving a minor coverage gap.

4 / 5

Completeness

It explicitly answers both questions: 'Use when a completed research manuscript needs a robust internal, venue-calibrated mock peer review' (when) plus the subagent orchestration and MetaReview duties (what), and adds scope boundaries ('Never revises research artifacts, executes experiments, or updates pipeline state'). It is not a 5 because the 'when' clause is a single trigger phrasing without the concrete trigger variations and synonyms the 5 anchor requires.

4 / 5

Trigger Term Quality

Natural trigger phrases like 'mock peer review', 'research manuscript', 'reviewer', and 'venue' are present and would be said by a real user. However, the most common user phrasing 'paper' / 'paper draft' and the venue names (ICLR, NeurIPS) that the body supports are absent, matching the 4 anchor ('good keyword coverage; a few natural terms missing') rather than the 5 anchor's comprehensive synonym coverage.

4 / 5

Distinctiveness Conflict Risk

It carves a clear niche — ensemble mock peer review of a manuscript with advisory reroutes — and the 'Never revises research artifacts, executes experiments, or updates pipeline state' clause actively separates it from overlapping paper-pipeline skills. Trigger terms ('mock peer review', 'manuscript', 'MetaReview') are distinct, so conflict risk is minimal.

5 / 5

Total

17

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

referenced_paths_exist

Referenced path issues: 6 deeper-than-1-level

Warning

Total

14

/

16

Passed

Repository
tsinghua-fib-lab/AgentSociety
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.