CtrlK
BlogDocsLog inGet started
Tessl Logo

strategy-pivot-designer

Detect backtest iteration stagnation and generate structurally different strategy pivot proposals when parameter tuning reaches a local optimum.

65

Quality

82%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

82%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-engineered skill body: concise, accurately executable commands whose flags match the real script CLIs, a clearly sequenced workflow with a decision gate and iteration feedback loop, and a sensible split between the overview and real one-level-deep reference files. Remaining improvements are small: replace wildcard path placeholders with concrete examples, add an output-validation step (schema check or bundled tests), and annotate the reference files with when-to-read guidance.

Suggestions

Use concrete file paths instead of wildcard placeholders in the generate_pivots example (the script's --diagnosis takes a single path, so a shell glob would silently expand to the wrong file if multiple diagnoses exist).

Add a validation checkpoint after pivot generation, e.g. verifying output YAML against references/pivot_proposal_schema.md or running the bundled scripts/tests, to strengthen the workflow's feedback loop.

Annotate the Resources entries (or add inline signposts in the Workflow section) so each reference file states when to consult it, moving navigation from adequate to fully signaled.

DimensionReasoningScore

Conciseness

The body is lean and assumes competence: every section (Overview, When to Use, Prerequisites, Output, Workflow, Quick Commands, Resources) is task-specific with zero filler, no explanation of concepts Claude already knows, and no padding. It matches the 'every token earns its place' anchor; nothing needed trimming to the level of 4.

5 / 5

Actionability

The three Quick Commands are concrete, copy-paste-ready bash invocations whose flags match the actual argparse interfaces of detect_stagnation.py and generate_pivots.py (verified). Minor gaps keep it below 5: the --diagnosis and --strategy arguments use wildcard placeholders (reports/pivot_diagnosis_*.json) that must be filled in — and the script accepts a single diagnosis path, so a shell glob would silently pass only the last match — and there is no example of interpreting the diagnosis or ranking output.

4 / 5

Workflow Clarity

The seven-step workflow is clearly sequenced, contains an explicit decision gate ("If stagnation detected, generate pivot proposals"), and closes a feedback loop by feeding selected pivots back into backtest-expert. It fits 'clear sequence with most checkpoints present; minor validation gaps': no step validates generated pivot YAML against the pivot proposal schema or mentions the bundled tests, and the file-generation steps produce batch outputs without a verify step. It is not capped at 3 because the operations are additive file generation, not destructive or destructive-batch changes.

4 / 5

Progressive Disclosure

The SKILL.md is a clean overview with details correctly split into four real, one-level-deep reference files (stagnation_triggers, strategy_archetypes, pivot_techniques, pivot_proposal_schema — all verified present with substantive content) and two scripts. It sits at 4 rather than 5 because the Resources section lists reference paths bare, without one-line annotations or in-text signposts (e.g., workflow step 3 names the three pivot techniques but never says 'see references/pivot_techniques.md'), so navigation is good but not fully signaled.

4 / 5

Total

17

/

20

Passed

Description

73%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, domain-specific description that clearly states what the skill does and gives an explicit when-clause. Its main limitation is coverage: only two of the skill's several capabilities and only one of the four stagnation triggers appear, which slightly undersells both the skill's actions and its trigger surface.

Suggestions

Extend the description with one or two more concrete actions (e.g., score and rank pivot proposals, emit exportable ticket YAML) to lift specificity toward comprehensive coverage.

Broaden the when-clause to include the other documented stagnation triggers (e.g., 'Use when backtest scores plateau, the strategy overfits, or transaction costs defeat the thin edge') to improve both trigger-term quality and completeness.

DimensionReasoningScore

Specificity

The description names exactly two concrete actions — "Detect backtest iteration stagnation and generate structurally different strategy pivot proposals" — which matches the '1-2 concrete actions, but not comprehensive' anchor. It omits real capabilities (ranked proposals, diagnosis reporting, exportable ticket generation) that would push it to 4, but it is well above the generic 'names the domain but actions are minimal' level of 2.

3 / 5

Completeness

It answers both parts: what ("Detect... and generate structurally different strategy pivot proposals") and when ("when parameter tuning reaches a local optimum"). The when-clause is explicit but narrow — it names only the local-optimum condition and misses the other stagnation triggers (overfitting, cost defeat, tail risk) the body documents, so it fits 'has both what and when; when could be more explicit or specific' rather than the fully concrete trigger coverage of 5.

4 / 5

Trigger Term Quality

Natural domain phrases like "backtest iteration stagnation", "parameter tuning", and "local optimum" are exactly what a user of this pipeline would say. A few natural synonyms from the skill's own When-to-Use list (plateau, overfitting, tail risk, transaction costs) are missing, which keeps it below the comprehensive synonym coverage of 5 but clearly above the sparse keyword level of 3.

4 / 5

Distinctiveness Conflict Risk

The vocabulary ("backtest iteration stagnation", "strategy pivot proposals", "parameter tuning reaches a local optimum") carves out a clear niche distinct from adjacent skills like strategy designers or parameter tuners, with minimal conflict risk. It is in no way generic, so the overlap-risk anchors at 1-4 do not apply.

5 / 5

Total

16

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
tradermonty/claude-trading-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.