Content
70%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
Operationally excellent content — executable commands, explicit validation gates, and regression tests — but it is a monolithic 695-line file that inlines substantial spec material that belongs in reference files, and it carries notable redundancy (duplicated reviewer prompt, rationale prose). Splitting opt-in sections (edit whitelist, style-ref, kill-argument, reviewer prompt template) into references/ would fix both the progressive-disclosure and conciseness weaknesses at once.
Suggestions
Move the ~100-line Edit Whitelist spec, the Style Reference contract, and the reviewer prompt template into separate files under references/ (e.g. references/edit-whitelist.md, references/reviewer-prompt.md) and keep one-line pointers plus defaults in SKILL.md — the prompts are duplicated verbatim between Step 2 and Step 5 and only the bias-guard delta needs to be inline.
Tighten or relocate the "Rationale" and "Empirical motivation" prose paragraphs (e.g. under Step 4.5, Step 5.5, Step 8) into a single 'why these rules exist' note or drop them; they justify design choices but cost context on every invocation.
Fix the duplicate step numbering (two 'Step 9' headings) and make the Step 4.5 restatement check fully executable — the inline Python currently stops at the normalize() function with the comparison left as a comment.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | Mostly skill-specific knowledge rather than concepts Claude already knows, but the ~50-line reviewer prompt is duplicated nearly verbatim in Steps 2 and 5, "Rationale"/"Empirical motivation" prose is padded, and the ~100-line edit-whitelist spec plus full prompt templates are inlined, so it could be noticeably tightened (not a 2 — the bulk is non-redundant operational detail). | 3 / 5 |
Actionability | Concrete, executable guidance dominates: exact bash commands (latexmk, pdfinfo, dup-label grep), full MCP call configs with model and reasoning-effort, a YAML whitelist schema, regex detectors, and invocation examples. Minor gaps: the Step 4.5 Python snippet ends in comments ("Compare normalized theorem blocks..." — the actual comparison logic is pseudocode), and [VENUE]/[list figure files] placeholders must be filled by hand. | 4 / 5 |
Workflow Clarity | Steps 0–9 are clearly sequenced with explicit validation checkpoints and feedback loops: recompile with "Verify: 0 undefined references, 0 undefined citations", the Step 4.5 restatement regression test rerun after every recompile, format-check stop criteria with severity thresholds, and auto-fix tables. Only slip is cosmetic: two steps are both labeled "Step 9" (Document Results / Summary), which is a labeling defect, not a missing validation. | 5 / 5 |
Progressive Disclosure | Section headers and in-body navigation are good, but the skill is a single 41KB monolithic file with no references/ bundle at all — the edit-whitelist spec, style-ref contract, restatement-check details, and both full reviewer prompts clearly belong in separate reference files ("content that should be separate is inline"). Cross-references (shared-references/*, tools/extract_paper_style.py, skills/kill-argument/SKILL.md) point outside the bundle and are at least clearly signaled, keeping this above the unstructured-inline anchor of 2. | 3 / 5 |
Total | 15 / 20 Passed |