Content
77%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
The body is a well-sequenced, actionable instruction skill with strong error handling and a calibrated-verdict design that resists false ABANDON/PROCEED errors. Its main weaknesses are dangling external references (none of the shared-references files or helper scripts exist in the bundle) and some duplication between the verdict limits and the Important Rules section.
Suggestions
Resolve the dangling references: either include `shared-references/integration-contract.md`, `citation-discipline.md`, `review-tracing.md`, `verify_papers.py`, and `save_trace.sh` in the bundle (e.g., under `references/` and `scripts/`) or inline the essential parts of each so no step depends on an unverifiable file.
De-duplicate the "Two failures waste months equally" guidance and the proximity/calibration rules that appear in both §The verdict limits and §Important Rules — state them once and cross-reference, freeing tokens for a concrete `verify_papers.py` invocation example.
Move the dense anti-hallucination/citation policy paragraph to a clearly signaled reference file (or tighten it to the 2–3 operative rules inline) so the main body stays a scannable overview.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | The body is efficient and assumes competence — no explanations of what arXiv or novelty means — with tight phase instructions ("Try at least 3 different query formulations per claim"). Minor trimming is possible: the "Important Rules" section restates the verdict limits ("Two failures waste months equally" appears in both §The verdict limits and §Important Rules), and the citation-policy paragraph is dense enough to belong in the referenced reference file. | 4 / 5 |
Actionability | Mostly executable guidance: a concrete MCP invocation block (model, config with `model_reasoning_effort: "xhigh"`, prompt template), a dossier-file pattern with exact contents, specific search targets ("ICLR 2025/2026, NeurIPS 2025", "year filters for 2024-2026"), and a full report template. Minor gaps keep it below 5: the MCP snippet leaves placeholders ("<absolute path to NOVELTY_DOSSIER.md>"), and required helpers (`verify_papers.py`, `save_trace.sh`) are referenced without any invocation example and do not resolve within this bundle. | 4 / 5 |
Workflow Clarity | Phases A–D are clearly sequenced (extract claims → multi-source search → cross-model verification → structured report) with explicit validation and feedback loops: pre-search paper verification, and Policy D1's degraded fallback ("if the helper is unresolved or its invocation fails, tag candidate entries [UNVERIFIED] and surface the uncertainty rather than dropping them"). Error recovery (fallback tagging, permissive-verdict tie-breaking) is spelled out, matching the top anchor. | 5 / 5 |
Progressive Disclosure | The in-file structure is well organized (Constants, Instructions by phase, Important Rules, Review Tracing), but every external reference — `../shared-references/integration-contract.md`, `../shared-references/citation-discipline.md`, `shared-references/review-tracing.md`, `verify_papers.py`, `save_trace.sh` — is dangling: no references/, scripts/, or shared-references directories exist in this bundle, so navigation is unverifiable and protocol detail that should live in those files is inlined in SKILL.md. This matches 'Some structure but could be better organized' rather than the well-placed/one-level-deep anchors. | 3 / 5 |
Total | 16 / 20 Passed |