Content
42%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
The body is well-sectioned and contains genuinely useful concrete guidance (authority order, matching tolerances, failure markers), but it is significantly undermined by an unresolved internal contradiction — the new canonical-source principle (DOI/CrossRef/arXiv first, Google Scholar demoted) coexists with old Google-Scholar-centric instructions in the example, Best Practices, and Summary — plus heavy cross-section duplication and a complete failure to link the existing references/ and scripts/ bundle files.
Suggestions
Resolve the Google Scholar contradiction: update the worked example, Best Practices, and Summary to fetch BibTeX from CrossRef/arXiv/publisher metadata per the Core Principle, or explicitly present the Google Scholar path as fallback-only the way Verification Principle 2 does.
Link the bundle from SKILL.md (e.g., "**Verification rules**: See references/verification-rules.md", "**API usage**: See references/api-usage.md", "**Common errors**: See references/common-errors.md") instead of inlining that detail, and deduplicate the twice-written failure-handling and Best Practices sections into one.
Delete the closing Summary section (or compress it to the Core Principle plus the [CITATION NEEDED] convention) — it restates the entire skill a third time.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | The ~200-line body is noticeably verbose with clear duplication: "Handling Verification Failures" appears as a full ## section and again nearly verbatim as a ### under Best Practices ("Don't guess", "[CITATION NEEDED]", "Notify the user"); the Best Practices items ("Never generate citations from memory", "Use WebSearch to find", "Verify promptly") restate the Verification Principles section; and the closing Summary re-explains the entire skill a third time. This matches anchor 2 ("Noticeably verbose; several unnecessary explanations or padded sections") rather than anchor 3, since multiple whole sections could be deleted with no information loss. | 2 / 5 |
Actionability | There is real concrete guidance — the preferred authority order (DOI → arXiv → CrossRef → Semantic Scholar → Zotero → Google Scholar), matching tolerances ("Year (±1 year difference allowed)"), a concrete failure marker ("[CITATION NEEDED]"), and example WebSearch queries ("Attention is All You Need Vaswani 2017"). But the guidance is undermined by direct contradictions: the Core Principle states "Google Scholar... is not the canonical verification authority", yet the example instructs "Click 'Cite' on Google Scholar → Select BibTeX format → Copy BibTeX entry" and Best Practices demand "Confirm on Google Scholar" / "Copy BibTeX from Google Scholar". As written, an agent cannot execute a single consistent procedure, which keeps this at anchor 3 ("Some concrete guidance but incomplete; missing key details") rather than anchor 4. | 3 / 5 |
Workflow Clarity | A clear ASCII workflow with a validation gate exists ("Add to bibliography" only after "Confirm paper details"), and failure handling forms a feedback loop (check spelling → try different queries → alternative sources → mark pending → notify user), so this is above anchors 1-2. It falls short of anchor 4 because two competing workflows coexist: the flow diagram says "Find DOI / arXiv ID / publisher page... Verify metadata with CrossRef / arXiv / Semantic Scholar", while the worked example and Summary route through Google Scholar ("Google Scholar to verify existence"). The batch/destructive cap at 3 does not apply (this is neither batch nor destructive), but the incoherent dual sequence prevents the clear single sequence anchor 4 requires. | 3 / 5 |
Progressive Disclosure | Scored against the actual bundle: references/ contains common-errors.md, verification-rules.md, and api-usage.md, and scripts/ contains verify-citations.py, api-clients.py, and format-checker.py — yet the SKILL.md body references none of them (zero links or path mentions), so the bundle is undiscoverable from the skill entry point. The body itself is well-sectioned with headers, so it is not the anchor-2 "minimal structure" / wall of text; but the ~200 lines inline detailed verification rules and error patterns that duplicate the reference files while leaving those files completely unsignaled. This matches anchor 3 ("references present but not clearly signaled; content that should be separate is inline") and rules out anchor 4 ("references mostly clear"). | 3 / 5 |
Total | 11 / 20 Passed |