CtrlK
BlogDocsLog inGet started
Tessl Logo

proof-writer

Writes rigorous mathematical proofs for ML/AI theory. Use when asked to prove a theorem, lemma, proposition, or corollary, fill in missing proof steps, formalize a proof sketch, 补全证明, 写证明, 证明某个命题, or determine whether a claimed proof can actually be completed under the stated assumptions.

72

Quality

90%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

81%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A rigorous, highly actionable instruction-only skill with an exemplary gated workflow and explicit honesty rules; its only real weaknesses are redundant restatement of the extraction inputs across three sections and a single-file layout that could offload detail into references.

Suggestions

Merge the extraction lists: keep the full 'Inputs' checklist once and have Step 1 reference it instead of restating claim/assumptions/notation/sketch, then drop the duplicated restatement in Step 2.

Move the three 'Output Modes' detail blocks and the proof-file template into a single reference file (e.g. references/output-formats.md), leaving a one-line pointer per mode in SKILL.md.

Trim near-duplicate guardrail statements (e.g. defaulting to PROOF_PACKAGE.md appears under Constants, Step 1, and Step 5) to a single authoritative statement.

DimensionReasoningScore

Conciseness

The body is directive and never explains concepts Claude already knows, but the extraction list is stated three times: "## Inputs" ("exact theorem / lemma / proposition / corollary statement ... notation ... proof sketch"), "Step 1: Extract: exact claim, assumptions, notation, proof sketch", and again in "Step 2: Normalize the Claim". This is the 'mostly efficient but could be tightened' anchor rather than anchor 4, where only minor instances of over-explanation remain. It is not anchor 2 because there is no concept-explanation padding.

3 / 5

Actionability

Guidance is fully concrete: an exact copy-paste output-file template ("## Required File Structure"), three enumerated output modes with their exact contents, named statuses, an explicit verification checklist, and precise formatting rules ("use $...$ for inline math"). Although code-free, the scoring note says instruction-only skills are not penalized for absence of code when guidance is actionable, and this guidance leaves nothing abstract.

5 / 5

Workflow Clarity

A clearly sequenced 6-step workflow with an explicit feasibility checkpoint (Step 3 triage), a final verification checklist (Step 6), and an error-recovery loop ("If a key step still cannot be justified, downgrade the status and write a blockage report instead of forcing a proof"). This matches the anchor-5 example of validation steps plus feedback loops.

5 / 5

Progressive Disclosure

The body is well-organized into clearly headed sections (Constants, Workflow, Required File Structure, Output Modes, Key Rules) in a single flat file, which matches 'good structure; most content is appropriately placed' with minor gaps. It does not reach anchor 5 because at ~220 lines the output-mode details and file-structure template could be split into one-level-deep reference files, and it is not anchor 3 since no content is mis-filed or hard to navigate.

4 / 5

Total

17

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A model description: third-person, concrete action list, explicit and detailed 'Use when' triggers, bilingual synonyms, and a well-bounded niche. It reads like the reference good examples and avoids every bad-example failure mode.

DimensionReasoningScore

Specificity

Multiple specific concrete actions are listed — "Writes rigorous mathematical proofs", "fill in missing proof steps", "formalize a proof sketch", "determine whether a claimed proof can actually be completed" — giving comprehensive coverage of the skill's capability surface in third-person voice. It clearly matches the 'multiple specific concrete actions; comprehensive coverage' anchor and exceeds the anchor-4 example, which has minor coverage gaps.

5 / 5

Completeness

It explicitly answers both what ("Writes rigorous mathematical proofs for ML/AI theory") and when ("Use when asked to prove a theorem ... or determine whether a claimed proof can actually be completed") with concrete trigger phrases, structurally identical to the anchor-5 example. The 'Use when' clause is explicit and detailed, so the anchor-4 'when could be more specific' cap does not apply.

5 / 5

Trigger Term Quality

Trigger coverage is unusually thorough: "prove a theorem, lemma, proposition, or corollary", "fill in missing proof steps", "formalize a proof sketch", plus Chinese synonyms (补全证明, 写证明, 证明某个命题) that mirror the English phrasings. This meets the 'comprehensive coverage including synonyms' anchor rather than the anchor-4 case of missing natural terms.

5 / 5

Distinctiveness Conflict Risk

It occupies a clear niche (mathematical proof drafting/verification for ML/AI theory) with distinct triggers around theorems, lemmas, and proof completion that few other skills would claim. Overlap with generic paper-editing or writing skills is minimal, matching the 'clear niche with distinct triggers; minimal conflict risk' anchor.

5 / 5

Total

20

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

14

/

16

Passed

Repository
wanshuiyin/Auto-claude-code-research-in-sleep
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.