CtrlK
BlogDocsLog inGet started
Tessl Logo

backtest-expert

Expert guidance for systematic backtesting of trading strategies. Use when developing, testing, stress-testing, or validating quantitative trading strategies. Covers "beating ideas to death" methodology, parameter robustness testing, slippage modeling, bias prevention, and interpreting backtest results. Applicable when user asks about backtesting, strategy validation, robustness testing, avoiding overfitting, or systematic trading development.

71

Quality

89%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

81%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, highly actionable skill body with a clear sequenced workflow, concrete thresholds, and a verified executable script. Its main weakness is redundancy — friction modeling and failure-pattern guidance are repeated both within the body and against the reference files.

Suggestions

Consolidate friction/slippage guidance into a single section: the 1.5-2x slippage, worst-case fills, and commission advice currently appears in Workflow steps 3-4, "Punish the Strategy", and "Critical Reminders".

Trim the "Common Failure Patterns" section to a one-line pointer to references/failed_tests.md, which already covers the patterns in detail, keeping only the inline red-flag reminders.

Cut or compress the "Discretionary vs Systematic Differences" section to 1-2 sentences, since its scope limitation is already implied by the description and the When to Use section.

DimensionReasoningScore

Conciseness

The body is mostly efficient and assumes domain competence, but friction/slippage guidance ("Increase slippage to 1.5-2x typical") is repeated across step 3, step 4, "Punish the Strategy", and "Critical Reminders", and the "Common Failure Patterns" section duplicates content already in references/failed_tests.md. It sits above anchor 2 (no padding with concepts Claude already knows) but clearly could be tightened.

3 / 5

Actionability

Guidance is fully concrete and executable: exact stress-test grids ("stop loss at 50%, 75%, 100%, 125%, 150% of baseline"), explicit sample-size thresholds ("30 trades... 100+... 200+"), and a copy-paste-ready script invocation whose flags match the actual script CLI.

5 / 5

Workflow Clarity

The six-step workflow (State Hypothesis → Codify → Initial Backtest → Stress Test → Out-of-Sample → Evaluate) is clearly sequenced with feedback loops ("If fundamentally broken, iterate on hypothesis"), explicit warning signs, and a Deploy/Refine/Abandon decision checkpoint at the end.

5 / 5

Progressive Disclosure

Structure is good: two one-level-deep references (methodology.md, failed_tests.md — both real files), each with "When to read" guidance and a contents list, plus a clearly documented script. Minor gap: failure-pattern content is duplicated inline in the body instead of being left to the reference file.

4 / 5

Total

17

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: it clearly states what the skill covers and when to use it, uses natural trigger phrases, and occupies a distinct niche. The only minor weakness is that a few common synonym trigger terms are absent.

DimensionReasoningScore

Specificity

The description lists multiple concrete capabilities — "parameter robustness testing, slippage modeling, bias prevention, and interpreting backtest results" — in third-person voice, giving comprehensive coverage of the skill's actions with no significant gaps.

5 / 5

Completeness

It explicitly answers both what ("Expert guidance for systematic backtesting... Covers 'beating ideas to death' methodology...") and when ("Use when developing, testing, stress-testing, or validating... Applicable when user asks about...") with concrete trigger phrases.

5 / 5

Trigger Term Quality

Trigger phrases like "backtesting, strategy validation, robustness testing, avoiding overfitting, or systematic trading development" are natural terms users would say, but a few common variations (e.g., curve-fitting, walk-forward analysis, "test my trading strategy") are missing, so it falls short of the synonym-complete anchor.

4 / 5

Distinctiveness Conflict Risk

The systematic-backtesting niche is well-defined with distinct triggers (backtesting, strategy validation, robustness testing), making conflict with other skills unlikely.

5 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
tradermonty/claude-trading-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.