CtrlK
BlogDocsLog inGet started
Tessl Logo

quant-statistics

Quantitative statistical methods: ADF unit-root / cointegration tests, GARCH volatility modeling, regression diagnostics (heteroskedasticity / autocorrelation), Bootstrap, and hypothesis testing.

60

Quality

76%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./agent/src/skills/quant-statistics/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

71%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A highly actionable reference skill: every method has executable code, exact return shapes, decision tables, and unusually good failure-mode documentation. The weaknesses are inlined textbook statistics that Claude already knows, and a monolithic single-file layout where domain heuristics and reference tables could be offloaded to bundle files to shrink the always-loaded surface.

Suggestions

Cut or compress the conceptual explanations Claude already knows (GARCH parameter meanings, the Bonferroni/Holm/FDR walkthrough, the hypothesis-testing quick-reference table, the Sharpe t-statistic derivation), keeping only the skill-specific guidance.

Move separable material — A-share/BTC parameter heuristics, the hypothesis-testing reference, and the output-format template — into one-level-deep reference files (e.g. references/domain-notes.md, references/output-format.md) and link them from SKILL.md.

Add a short end-to-end worked example (stationarity check → cointegration → hedge ratio → GARCH → bootstrap CI) that sequences the existing pieces into one validated flow.

DimensionReasoningScore

Conciseness

The API usage, return-dict shapes, error semantics, and measured gotchas (13% white-noise flag rate at lags=10, the ddof=1 vs ddof=0 denominator difference) earn their tokens, but several inlined passages explain concepts Claude already knows — GARCH parameter meanings, Bonferroni/Holm/Benjamini-Hochberg corrections, the hypothesis-testing quick-reference table, and the Sharpe t-statistic formula. Not 4: more than minor instances of over-explanation could be trimmed; not 2: the bulk is skill-specific rather than padded.

3 / 5

Actionability

Every section leads with copy-paste-ready imports and calls showing exact signatures, real return dicts, and decision tables (p-value bands to actions, z-score bands to signals), plus precise input conventions like "returns are FRACTIONS (0.01 = 1%)". Not 4: the common cases are fully executable with concrete expected outputs.

5 / 5

Workflow Clarity

The "import and call it — do not retype" rule, ordered prerequisites (run adf_test on both legs before cointegration), a 6-item regression diagnostics checklist, and error-recovery guidance (ImportError → report to the user rather than silently substituting) provide a clear, checkpointed flow. Not 5: there is no explicit end-to-end sequence with validate-and-retry loops; not 3: checkpoints are present via the checklist, preconditions, and degenerate-input handling.

4 / 5

Progressive Disclosure

No bundle files exist, and all ~370 lines are inlined in SKILL.md. Section headers are clear, but separable material — the China A-share/BTC parameter lore, the hypothesis-testing quick reference, and the output-format template — stays in the main file instead of being split into one-level-deep reference files. Not 4: content that should live in separate files is inline; not 2: structure is strong with well-organized, navigable sections.

3 / 5

Total

15

/

20

Passed

Description

71%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A specific, well-enumerated description that clearly communicates what the skill does and is largely distinct from neighboring skills. Its main weakness is the complete absence of a "when to use" clause, which both caps completeness and leaves natural trigger phrasings like stationarity or volatility forecasting uncovered.

Suggestions

Append an explicit trigger clause, e.g. "Use when checking stationarity, testing cointegration for pair trading, forecasting volatility, or assessing statistical significance of backtests and factor returns."

Add natural user phrasings and synonyms (stationarity, volatility forecasting, p-values, statistical significance, pair trading) so the description matches how users actually ask for these tasks.

Consider trimming the generic tail terms ("Bootstrap, and hypothesis testing") in favor of the more distinctive named tests to further reduce overlap with general statistics skills.

DimensionReasoningScore

Specificity

Lists multiple specific concrete actions — "ADF unit-root / cointegration tests", "GARCH volatility modeling", "regression diagnostics (heteroskedasticity / autocorrelation)", "Bootstrap, and hypothesis testing" — comprehensively covering the skill's scope. Not 4: there are no coverage gaps; the enumeration matches the body's actual sections.

5 / 5

Completeness

The "what" is explicit and concrete, but there is no "Use when..." clause or any equivalent explicit trigger guidance, which caps completeness at 3. Not 4: the "when" is entirely missing, not merely implicit or under-specified.

3 / 5

Trigger Term Quality

Good keyword coverage with domain terms users would naturally say (cointegration, GARCH, heteroskedasticity, autocorrelation, bootstrap, hypothesis testing), but common variations and synonyms are missing (stationarity, volatility forecasting, p-value / statistical significance, pair trading). Not 5: the synonym layer of anchor 5 is absent.

4 / 5

Distinctiveness Conflict Risk

Named tests (ADF, cointegration, GARCH) carve a distinct quant-statistics niche with minimal conflict risk, though the generic tail terms "Bootstrap, and hypothesis testing" leave minor overlap with a general statistics skill. Not 5: those broader terms could still match non-quant requests.

4 / 5

Total

16

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
HKUDS/Vibe-Trading
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.