CtrlK
BlogDocsLog inGet started
Tessl Logo

proof-checker

Rigorous mathematical proof verification and fixing workflow. Reads a LaTeX proof, identifies gaps via cross-model review (external reviewer backend, ultra reasoning), fixes each gap with full derivations, re-reviews, and generates an audit report. Use when user says "检查证明", "verify proof", "proof check", "审证明", "check this proof", or wants rigorous mathematical verification of a theory paper.

70

Quality

88%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

77%

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A highly actionable, well-sequenced workflow with concrete prompts, tool calls, and validation checkpoints. Its main weaknesses are verbosity from repeating opt-in/non-blocking invariants and broken/missing shared-references bundle files that the body relies on.

Suggestions

Consolidate the repeated 'absent == unavailable == not blocking' explanations for deep-fix and restatement-check into a single canonical statement in Submission Artifact Emission, then reference it from Key Rules and the phase sections to remove the duplicated prose.

Create the referenced ../shared-references/*.md files (reviewer-routing.md, fan-out-pattern.md, assurance-contract.md, acceptance-gate.md, external-cadence.md, reviewer-independence.md) or remove the dangling links, since progressive disclosure depends on them resolving.

Move the long deep-fix prompt block, restatement-check algorithm, and JSON schema examples into dedicated reference files so the SKILL.md body stays a lean overview.

DimensionReasoningScore

Conciseness

Mostly domain-specific material Claude would not already know, but several invariants are re-explained multiple times — e.g., the 'absent == unavailable == not blocking' deep-fix/restatement semantics repeat across Phase 1, the Deep-Fix Mode section, Key Rules, and Submission Artifact Emission. It is accurate but could be tightened.

2 / 3

Actionability

Provides exact reviewer prompts, exact MCP tool calls with configs, exact bash/compile commands, fix-record templates, and JSON schemas — copy-paste ready and fully executable, matching the score-3 anchor.

3 / 3

Workflow Clarity

Clear phased sequence (Phase 0 through 5.5) with explicit validation checkpoints (acceptance gate, compile checks, blind re-review, regression audit) and feedback loops (repeat Phases 2-3 up to MAX_REVIEW_ROUNDS; re-enter Phase 2 on new issues), appropriate for batch .tex edits.

3 / 3

Progressive Disclosure

References to ../shared-references/*.md (reviewer-routing, fan-out-pattern, assurance-contract, etc.) are clearly signaled and one-level-deep by design, but none of those files exist in the workspace (dangling links), and large blocks like the deep-fix prompt, restatement algorithm, and JSON schemas are inline rather than split into the referenced files.

2 / 3

Total

10

/

12

Passed

Description

100%

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: third-person voice, concrete actions, explicit multilingual triggers, and a clear what+when structure. It distinguishes itself well from other skills and would reliably surface for the intended use cases.

DimensionReasoningScore

Specificity

Lists multiple concrete actions: 'Reads a LaTeX proof, identifies gaps via cross-model review', 'fixes each gap with full derivations, re-reviews, and generates an audit report' — a comprehensive action set matching the score-3 anchor.

3 / 3

Completeness

Clearly answers both what (verification+fixing workflow with concrete steps) and when via an explicit 'Use when user says...' clause, matching the score-3 anchor exactly.

3 / 3

Trigger Term Quality

Natural multilingual triggers users would actually say ('verify proof', 'proof check', 'check this proof', '检查证明', '审证明') plus the descriptive trigger 'rigorous mathematical verification of a theory paper', giving good coverage.

3 / 3

Distinctiveness Conflict Risk

Niche is precise (rigorous mathematical proof verification of theory papers) with distinct trigger phrases unlikely to fire for unrelated skills.

3 / 3

Total

12

/

12

Passed

Validation

75%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation12 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (860 lines); consider splitting into references/ and linking

Warning

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

relative_links

Relative link issues: 4 suspicious

Warning

Total

12

/

16

Passed

Repository
wanshuiyin/Auto-claude-code-research-in-sleep
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.