CtrlK
BlogDocsLog inGet started
Tessl Logo

math-olympiad

Solve competition math problems (IMO, Putnam, USAMO, AIME) with adversarial verification that catches the errors self-verification misses. Activates when asked to 'solve this IMO problem', 'prove this olympiad inequality', 'verify this competition proof', 'find a counterexample', 'is this proof correct', or for any problem with 'IMO', 'Putnam', 'USAMO', 'olympiad', or 'competition math' in it. Uses pure reasoning (no tools) — then a fresh-context adversarial verifier attacks the proof using specific failure patterns, not generic 'check logic'. Outputs calibrated confidence — will say 'no confident solution' rather than bluff. If LaTeX is available, produces a clean PDF after verification passes.

76

Quality

96%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

92%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A highly actionable, well-sequenced workflow body with strong progressive disclosure and explicit validation feedback loops. The only weakness is mild redundancy between the inline workflow and the closing recap section, which inflates token cost without adding new guidance.

Suggestions

Trim or fold the closing 'What makes this different from generic verify-and-refine' section — its six points (dual context isolation, pattern-specific attacks, asymmetric vote, spec-gaming, calibrated abstention, presentation pass) are already detailed in steps 1-8.

State the 50/63 interpretation-trap statistic once (step 1) rather than repeating the framing in both the intro list and the closing section.

Move the per-model solver/verify-pass table out of SKILL.md into references/model_tier_defaults.md (already referenced) and keep only a one-line pointer, since the body already directs readers there.

DimensionReasoningScore

Conciseness

Mostly efficient and assumes Claude's intelligence (no explanations of what an olympiad is), but the closing 'What makes this different' section and the 50/63 interpretation anecdote repeat points already made in steps 1, 3, and 5 and could be trimmed.

4 / 5

Actionability

Copy-paste-ready verbatim solver/verifier/reviser prompts, concrete agent counts (8-12 attempts, up to 5 verifiers), explicit return-format templates, and real script invocations (scripts/check_latex.sh, scripts/compile_pdf.sh) cover the common cases.

5 / 5

Workflow Clarity

A clearly numbered 8-step sequence with explicit validation checkpoints (step 4 adversarial verify, step 5 vote-verify) and feedback loops (hole found → revise → re-vote), including batch/parallel handling with the opts.label problem-ID discipline.

5 / 5

Progressive Disclosure

SKILL.md is a well-signaled overview with one-level-deep references (references/adversarial_prompts.md, verifier_patterns.md, presentation_prompts.md, model_tier_defaults.md, solver_heuristics.md) — all of which resolve to real files — and a closing 'Key references' index for easy navigation.

5 / 5

Total

19

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An excellent description: third-person voice, concrete actions, explicit 'use when' trigger guidance with natural user phrasings, and a well-scoped niche with low conflict risk. It satisfies every anchor at the top of the scale without padding.

DimensionReasoningScore

Specificity

Lists multiple specific concrete actions — 'Solve competition math problems', 'a fresh-context adversarial verifier attacks the proof using specific failure patterns', 'Outputs calibrated confidence', 'produces a clean PDF' — with comprehensive coverage rather than generic language.

5 / 5

Completeness

Clearly answers both what ('Solve competition math problems ... with adversarial verification') and when ('Activates when asked to ... or for any problem with IMO, Putnam, USAMO, olympiad, or competition math in it') with concrete trigger phrases.

5 / 5

Trigger Term Quality

Comprehensive natural trigger phrases users would actually say ('solve this IMO problem', 'prove this olympiad inequality', 'verify this competition proof', 'find a counterexample', 'is this proof correct') plus synonyms and named contests (IMO, Putnam, USAMO, AIME, olympiad, competition math).

5 / 5

Distinctiveness Conflict Risk

A clear niche — verified competition math on named contests — with distinct triggers; the adversarial-verification framing and named olympiads make overlap with other skills minimal.

5 / 5

Total

20

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
anthropics/claude-plugins-official
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.