CtrlK
BlogDocsLog inGet started
Tessl Logo

paper-claim-audit

Zero-context verification that every number, comparison, and scope claim in the paper matches raw result files. Uses a fresh cross-model reviewer with NO prior context to prevent confirmation bias. Use when user says "审查论文数据", "check paper claims", "verify numbers", "论文数字核对", or before submission to ensure paper-to-evidence fidelity.

70

Quality

88%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

81%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A highly actionable, well-sequenced audit skill with concrete cross-model reviewer prompts, report templates, and a verdict decision table. Its main weakness is redundancy: the zero-context fresh-reviewer invariant is restated several times and some motivational sections could be trimmed.

Suggestions

Consolidate the zero-context / fresh-thread / never-codex-reply rule into a single canonical statement and reference it rather than restating it in Core Principle, Step 2, Thread independence, and Key Rules.

Trim or fold the "Why This Exists" and "How This Differs From Other Audit Skills" sections into a brief motivating paragraph; the comparison table and bias examples are explanatory padding that competes with the operational context budget.

Consider moving the full audited_input_hashes JSON schema and verdict decision table into a reference file under a references/ bundle, keeping SKILL.md as an overview with the essential codex prompt and report skeleton inline.

DimensionReasoningScore

Conciseness

The zero-context / fresh-thread / never-codex-reply rule is repeated across Core Principle, Step 2's CRITICAL note, the Thread independence section, and Key Rules, and the "Why This Exists" / "How This Differs" sections are motivational padding; the operational core is efficient but the file could be tightened.

3 / 5

Actionability

Provides a copy-paste-ready codex invocation with explicit model and reasoning-effort config, a full audit protocol with seven named failure modes, and concrete report templates for both Markdown and JSON output, covering the common cases.

5 / 5

Workflow Clarity

A clear four-step sequence (Collect Files → Fresh Reviewer Audit → Write Report → Print Summary) with explicit validation via the verdict decision table and status enum, plus a feedback loop into the improvement round; the skill is read-only so the batch/destructive validation cap does not apply.

5 / 5

Progressive Disclosure

Content is well-organized into clearly headed sections with one-level-deep, clearly signaled references to external shared-references files (external-cadence.md, review-tracing.md, assurance-contract.md, reviewer-independence.md); no bundle directory exists, and the large inlined codex prompt and schema are core operational content that belongs inline, leaving only minor organization gaps.

4 / 5

Total

17

/

20

Passed

Description

95%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, well-structured description that clearly states both capability and trigger conditions with bilingual natural phrases. The only minor gap is that it names one verifying action across multiple claim types rather than several distinct operations.

DimensionReasoningScore

Specificity

Describes a concrete verification action ("every number, comparison, and scope claim in the paper matches raw result files") and the fresh cross-model reviewer mechanism, but it is one verify action applied to several claim types rather than multiple distinct concrete actions, so it stops short of the comprehensive anchor 5.

4 / 5

Completeness

Explicitly answers both what ("Zero-context verification that every number, comparison, and scope claim in the paper matches raw result files") and when ("Use when user says ... or before submission"), with concrete trigger phrases as in the anchor 5 example.

5 / 5

Trigger Term Quality

Provides comprehensive natural trigger phrases with synonyms in both English and Chinese ("check paper claims", "verify numbers", "审查论文数据", "论文数字核对") plus the "before submission" usage cue, matching the comprehensive-coverage anchor.

5 / 5

Distinctiveness Conflict Risk

The paper-to-evidence fidelity framing with a zero-context fresh reviewer carves a clear niche distinct from general fact-checking or experiment-audit skills, with minimal conflict risk.

5 / 5

Total

19

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

relative_links

Relative link issues: 1 suspicious

Warning

Total

13

/

16

Passed

Repository
wanshuiyin/Auto-claude-code-research-in-sleep
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.