CtrlK
BlogDocsLog inGet started
Tessl Logo

statistical-reporting

Statistical test selection, assumption checking, and APA-formatted reporting. Use when analyzing experimental results or writing results sections.

68

Quality

82%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

The canonical home for this skill is statistical-reporting in aiming-lab/AutoResearchClaw

SKILL.md
Quality
Evals
Security

Quality

Content

86%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a tight, well-structured statistical-reporting reference that adds genuinely non-obvious detail (exact APA formats, effect-size conventions, common pitfalls) without padding. Its main limitation is that it is a lookup reference rather than a runnable worked workflow.

Suggestions

Add one short worked example (a tiny dataset → chosen test → full APA-reported sentence) to move actionability from concrete templates toward copy-paste-ready output.

Consider an explicit 3-4 step selection workflow (e.g., '1. Check assumptions, 2. Pick test from table, 3. Run, 4. Report with effect size') to give workflow_clarity a clearer sequence.

Optionally split the per-test APA format strings into a short references table if the skill grows, keeping the main body as an overview.

DimensionReasoningScore

Conciseness

The body is a lean reference with no padding — it never explains what a t-test or ANOVA is, and every line states a test, threshold, or format that Claude would not reliably reproduce from memory, matching the 'lean and efficient; every token earns its place' anchor; it is clearly above the 'minor over-explanation' (4) anchor.

5 / 5

Actionability

Provides concrete, copy-ready guidance — exact APA templates ('t(df) = X.XX, p = .XXX, d = X.XX') and specific thresholds (VIF < 5, Cohen's d 0.2/0.5/0.8) — but stops at lookup tables rather than giving a runnable worked example or sample analysis call, fitting 'mostly executable guidance; minor gaps' and falling short of the fully copy-paste-ready (5) anchor.

4 / 5

Workflow Clarity

The content is organized into a coherent decision framework (test selection list, then assumptions, then reporting format, then effect sizes, then pitfalls) but presents static reference rather than an explicit step sequence; since this is not a destructive or batch operation the 3-cap does not apply, and the clear organization places it above 'sequence present but checkpoints missing' (3) yet below the explicit validate-retry (5) anchor.

4 / 5

Progressive Disclosure

The body is under 50 lines, self-contained with no external references, and divided into well-labeled sections (Test Selection, Assumption Checking, APA Reporting, Effect Sizes, Common Mistakes); per the simple-skill exception, this warrants the top anchor without separate reference files.

5 / 5

Total

18

/

20

Passed

Description

78%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is concise, third-person, and explicitly pairs a concrete 'what' with a 'Use when' trigger clause. It is specific and distinct, with only minor gaps in trigger-phrase breadth and a small overlap risk with general writing skills.

Suggestions

Broaden the 'Use when' clause with concrete trigger keywords users actually say (e.g., 'hypothesis testing', 'p-values', 'ANOVA', 'regression results', 'effect sizes') to raise trigger-term coverage toward 5.

Add one more concrete capability (e.g., 'power analysis' or 'diagnostic assumption checks') to push specificity from several actions toward comprehensive coverage.

Sharpen distinctiveness by signaling the statistical/quantitative focus earlier in the sentence to reduce overlap with generic scientific-writing skills.

DimensionReasoningScore

Specificity

Lists three concrete actions — 'Statistical test selection, assumption checking, and APA-formatted reporting' — which is more than the 1-2 of a score-3 anchor, though coverage is not exhaustive (e.g., no mention of diagnostic plotting or power analysis), fitting the 'several specific actions; minor gaps' anchor.

4 / 5

Completeness

Explicitly answers both what ('Statistical test selection, assumption checking, and APA-formatted reporting') and when ('Use when analyzing experimental results or writing results sections'); the 'when' is explicit but offers only two trigger scenarios rather than the concrete, varied trigger phrases of a score-5 anchor, so it fits the 'both present, when could be more explicit' anchor.

4 / 5

Trigger Term Quality

Includes natural phrases users would say — 'analyzing experimental results', 'writing results sections', 'statistical test' — but omits common synonyms like 'hypothesis test', 'p-value', or 'regression' from the description itself, so it falls short of comprehensive (5) yet sits clearly above 'some relevant keywords' (3).

4 / 5

Distinctiveness Conflict Risk

The combination of statistical test selection and APA-formatted reporting carves a clear niche with distinct triggers, but 'writing results sections' creates minor overlap risk with general scientific-writing skills, matching 'mostly distinct; minor overlap risk' rather than the minimal-conflict score-5 anchor.

4 / 5

Total

16

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
aiming-lab/AutoResearchClaw
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.