CtrlK
BlogDocsLog inGet started
Tessl Logo

backtest-diagnose

Diagnose failed or underperforming backtests, locate the root cause, and fix the issue

62

Quality

78%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./agent/src/skills/backtest-diagnose/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

81%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a strong, lean diagnostic playbook: a sequenced workflow with explicit validation checkpoints, a concrete error taxonomy, hard gates, and guardrails (ignore list, iteration cap, precise-edit principle). Minor gains are available from deduplicating the zero-trades logic and tightening a few high-level fix hints.

DimensionReasoningScore

Conciseness

The body is dense and domain-specific — error taxonomy tables, a hard-gate checklist, and fix principles — with no padding explaining concepts Claude already knows. Minor instances could be trimmed (the zero-trades case appears in both the Logic Bugs section and the Hard-Gate Checklist), so it sits just below the lean anchor 5.

4 / 5

Actionability

Concrete file paths (artifacts/metrics.csv, code/signal_engine.py), a copy-paste AST syntax check command, and specific action-item examples ("Change RSI threshold from 30 to 25 in signal_engine.py line 42") make the guidance mostly executable. Some taxonomy fixes remain high-level hints ("Add length checks", "Check the actual column names in data_map") and `pip install xxx` is a placeholder pattern, keeping it below anchor 5.

4 / 5

Workflow Clarity

The 5-step Diagnostic Workflow is clearly sequenced with an explicit "Verify the fix" step, backed by the Hard-Gate Checklist and Post-Fix Validation Rules (AST parse, structural assertions, rerun). Feedback loops are explicit — rerun immediately after each fix, verify new metrics — with an iteration cap of 3, matching the anchor for clear sequence with explicit validation and error recovery.

5 / 5

Progressive Disclosure

No bundle files exist and the single SKILL.md is well organized into clearly headed sections with no inlined content that belongs in separate files; manifest-format details are correctly deferred to the external `strategy-discovery` skill via a clear one-level reference. It falls just short of anchor 5's ideal of well-signaled references throughout, since the Evidence hookup reference points to another skill rather than explicit reference files and the ~93-line body has no progressive-disclosure structure beyond headings.

4 / 5

Total

17

/

20

Passed

Description

66%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is concise, third-person, and domain-specific with a clear statement of what the skill does, but it lacks an explicit "Use when..." trigger clause, which both caps completeness and leaves natural trigger phrases (error, poor results) unused. Overall a solid but incomplete description.

Suggestions

Add an explicit trigger clause, e.g. "Use when a backtest fails, raises an error, exits nonzero, or produces poor results" — this raises completeness from 3 to 4-5 and mirrors the triggers already stated in the body's Overview.

Include the natural synonym phrases users actually report ("error", "poor results", "trading strategy") so trigger matching catches phrasings beyond "failed/underperforming backtests".

Optionally mention a distinctive capability such as artifact-based diagnosis (inspecting metrics/equity/trades CSVs) to differentiate from generic debugging skills and sharpen specificity.

DimensionReasoningScore

Specificity

"Diagnose failed or underperforming backtests, locate the root cause, and fix the issue" names the domain and three concrete actions (diagnose, locate, fix). Not 3 because it lists several specific actions rather than just 1-2; not 5 because diagnose/locate/fix are near-synonymous stems and coverage gaps exist (no mention of artifact inspection, error classification, or validation).

4 / 5

Completeness

The "what" is clear (diagnose, locate root cause, fix), but there is no "Use when..." clause or equivalent explicit trigger guidance in the description — the trigger conditions live only in the body's Overview ("when a user reports that a backtest failed, raised an error, or produced poor results"). Per the judging guidelines, a missing explicit trigger clause caps completeness at 3.

3 / 5

Trigger Term Quality

"backtests", "failed", "underperforming", "fix" are natural phrases a user would say when needing this skill. A few common variations are missing ("error", "poor results", "losing money"), which the body's Overview mentions but the description omits, so it falls short of comprehensive anchor 5.

4 / 5

Distinctiveness Conflict Risk

"backtest" is a distinct niche trigger unlikely to fire for unrelated skills, with minimal overlap risk against general debugging or data-analysis skills. Generic tail phrases like "fix the issue" carry minor overlap risk, keeping it below anchor 5.

4 / 5

Total

15

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
HKUDS/Vibe-Trading
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.