Content
71%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
A highly actionable reference skill: every method has executable code, exact return shapes, decision tables, and unusually good failure-mode documentation. The weaknesses are inlined textbook statistics that Claude already knows, and a monolithic single-file layout where domain heuristics and reference tables could be offloaded to bundle files to shrink the always-loaded surface.
Suggestions
Cut or compress the conceptual explanations Claude already knows (GARCH parameter meanings, the Bonferroni/Holm/FDR walkthrough, the hypothesis-testing quick-reference table, the Sharpe t-statistic derivation), keeping only the skill-specific guidance.
Move separable material — A-share/BTC parameter heuristics, the hypothesis-testing reference, and the output-format template — into one-level-deep reference files (e.g. references/domain-notes.md, references/output-format.md) and link them from SKILL.md.
Add a short end-to-end worked example (stationarity check → cointegration → hedge ratio → GARCH → bootstrap CI) that sequences the existing pieces into one validated flow.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | The API usage, return-dict shapes, error semantics, and measured gotchas (13% white-noise flag rate at lags=10, the ddof=1 vs ddof=0 denominator difference) earn their tokens, but several inlined passages explain concepts Claude already knows — GARCH parameter meanings, Bonferroni/Holm/Benjamini-Hochberg corrections, the hypothesis-testing quick-reference table, and the Sharpe t-statistic formula. Not 4: more than minor instances of over-explanation could be trimmed; not 2: the bulk is skill-specific rather than padded. | 3 / 5 |
Actionability | Every section leads with copy-paste-ready imports and calls showing exact signatures, real return dicts, and decision tables (p-value bands to actions, z-score bands to signals), plus precise input conventions like "returns are FRACTIONS (0.01 = 1%)". Not 4: the common cases are fully executable with concrete expected outputs. | 5 / 5 |
Workflow Clarity | The "import and call it — do not retype" rule, ordered prerequisites (run adf_test on both legs before cointegration), a 6-item regression diagnostics checklist, and error-recovery guidance (ImportError → report to the user rather than silently substituting) provide a clear, checkpointed flow. Not 5: there is no explicit end-to-end sequence with validate-and-retry loops; not 3: checkpoints are present via the checklist, preconditions, and degenerate-input handling. | 4 / 5 |
Progressive Disclosure | No bundle files exist, and all ~370 lines are inlined in SKILL.md. Section headers are clear, but separable material — the China A-share/BTC parameter lore, the hypothesis-testing quick reference, and the output-format template — stays in the main file instead of being split into one-level-deep reference files. Not 4: content that should live in separate files is inline; not 2: structure is strong with well-organized, navigable sections. | 3 / 5 |
Total | 15 / 20 Passed |