CtrlK
BlogDocsLog inGet started
Tessl Logo

statistical-analysis

Apply statistical methods including descriptive stats, trend analysis, outlier detection, and hypothesis testing. Use when analyzing distributions, testing for significance, detecting anomalies, computing correlations, or interpreting statistical results.

64

Quality

81%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

63%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is well-organized and actionably written, with executable outlier-detection code, a practical test-selection table, and business-oriented reporting guidance. Its main weaknesses are re-teaching statistics Claude already knows (hurting token efficiency) and keeping everything in one long file with no progressive disclosure via reference files.

Suggestions

Trim or drop the definitional explanations of concepts Claude already knows (what standard deviation/IQR are, the null-hypothesis framework, alpha=0.05) and keep only the applied decision guidance — e.g., replace the hypothesis-testing framework walkthrough with just the test-selection table and interpretation rules.

Split the self-contained sections into one-level-deep reference files (e.g., references/outlier-detection.md, references/hypothesis-testing.md, references/statistical-pitfalls.md) and keep SKILL.md as a concise overview with clearly signaled links.

Add small runnable snippets for the trend/forecasting guidance (e.g., a 3-line seasonal-naive or linear-trend forecast in pandas) so those sections are as executable as the outlier-detection code.

DimensionReasoningScore

Conciseness

The body spends tokens re-teaching textbook statistics Claude already knows ("Standard deviation: How far values typically fall from the mean", "Null hypothesis (H0): There is no difference", "Choose significance level (alpha): Typically 0.05"), alongside genuinely valuable applied guidance ("Always report mean and median together for business metrics"). This fits 'mostly efficient but includes some unnecessary explanation or could be tightened' — not level 2 because the majority is application/reporting guidance rather than padding.

3 / 5

Actionability

Concrete, executable Python snippets are provided for outlier detection (z-score, IQR, percentile) and moving averages ("df['ma_7d'] = df['metric'].rolling(window=7, min_periods=1).mean()"), plus a concrete test-selection table. Anchor 4 fits: mostly executable with minor gaps, since the forecasting and hypothesis-testing sections give prose guidance rather than runnable code.

4 / 5

Workflow Clarity

Multi-step processes are clearly sequenced ("1. Compute expected value... 4. Distinguish between point anomalies and change points"; the outlier handling decision tree "Investigate: Is this a data error, a genuine extreme value, or a different population?") with reporting checkpoints ("Report what you did"). Not level 5 because there are no explicit validate-and-recover feedback loops, though nothing destructive/batch applies the level-3 cap.

4 / 5

Progressive Disclosure

The skill has no bundle files — all ~245 lines live inline in SKILL.md with good section headers. It exceeds the <50-line simple-skill exception, and substantial content (hypothesis testing basics, the caution section on statistical claims) clearly belongs in separate reference files, matching anchor 3: 'some structure... content that should be separate is inline'.

3 / 5

Total

14

/

20

Passed

Description

88%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description that explicitly names its capabilities and pairs them with a concrete 'Use when...' trigger clause in third-person imperative voice. Keyword coverage is good but would benefit from common synonyms like 'A/B test' or 'p-value' to reach the top level.

DimensionReasoningScore

Specificity

The description lists multiple specific concrete capabilities — "descriptive stats, trend analysis, outlier detection, and hypothesis testing" — that comprehensively mirror the skill body's four sections, matching the anchor for comprehensive coverage rather than the 'minor gaps' level below.

5 / 5

Completeness

It clearly answers both questions: the 'what' ("Apply statistical methods including descriptive stats, trend analysis, outlier detection, and hypothesis testing") and an explicit 'Use when...' clause with concrete trigger phrases, exactly matching the top anchor.

5 / 5

Trigger Term Quality

Triggers like "analyzing distributions, testing for significance, detecting anomalies, computing correlations" are natural user phrasings, but common variations such as "A/B test", "p-value", or "statistically significant" are absent, so coverage is good but not comprehensive.

4 / 5

Distinctiveness Conflict Risk

The statistical-analysis niche is clear and mostly distinct, but broad phrases like "computing correlations" and "analyzing distributions" carry minor overlap risk with a general data-analysis skill, fitting the 'mostly distinct; minor overlap risk' anchor.

4 / 5

Total

18

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
anthropics/knowledge-work-plugins
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.