Content
81%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
A well-structured, highly actionable skill body with a clear sequenced workflow, concrete thresholds, and a verified executable script. Its main weakness is redundancy — friction modeling and failure-pattern guidance are repeated both within the body and against the reference files.
Suggestions
Consolidate friction/slippage guidance into a single section: the 1.5-2x slippage, worst-case fills, and commission advice currently appears in Workflow steps 3-4, "Punish the Strategy", and "Critical Reminders".
Trim the "Common Failure Patterns" section to a one-line pointer to references/failed_tests.md, which already covers the patterns in detail, keeping only the inline red-flag reminders.
Cut or compress the "Discretionary vs Systematic Differences" section to 1-2 sentences, since its scope limitation is already implied by the description and the When to Use section.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | The body is mostly efficient and assumes domain competence, but friction/slippage guidance ("Increase slippage to 1.5-2x typical") is repeated across step 3, step 4, "Punish the Strategy", and "Critical Reminders", and the "Common Failure Patterns" section duplicates content already in references/failed_tests.md. It sits above anchor 2 (no padding with concepts Claude already knows) but clearly could be tightened. | 3 / 5 |
Actionability | Guidance is fully concrete and executable: exact stress-test grids ("stop loss at 50%, 75%, 100%, 125%, 150% of baseline"), explicit sample-size thresholds ("30 trades... 100+... 200+"), and a copy-paste-ready script invocation whose flags match the actual script CLI. | 5 / 5 |
Workflow Clarity | The six-step workflow (State Hypothesis → Codify → Initial Backtest → Stress Test → Out-of-Sample → Evaluate) is clearly sequenced with feedback loops ("If fundamentally broken, iterate on hypothesis"), explicit warning signs, and a Deploy/Refine/Abandon decision checkpoint at the end. | 5 / 5 |
Progressive Disclosure | Structure is good: two one-level-deep references (methodology.md, failed_tests.md — both real files), each with "When to read" guidance and a contents list, plus a clearly documented script. Minor gap: failure-pattern content is duplicated inline in the body instead of being left to the reference file. | 4 / 5 |
Total | 17 / 20 Passed |