Guided statistical analysis for test selection, assumption checks, power analysis, and APA-style reporting. Use when you need to choose an appropriate statistical test for your data and produce publication-ready results (including effect sizes and diagnostics).
68
83%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
View guide
Passed
No findings from the security scan
Use this skill when you need to:
For programming model-specific workflows (especially regression variants and custom diagnostics), prefer statsmodels directly; this skill focuses on guided selection, checks, interpretation, and reporting.
references/test_selection_guide.md).scripts/assumption_checks.py (see references/assumptions_and_diagnostics.md).references/bayesian_statistics.md).references/effect_sizes_and_power.md).references/reporting_standards.md).Python (recommended 3.10+) with:
numpy>=1.24pandas>=2.0scipy>=1.10statsmodels>=0.14pingouin>=0.5matplotlib>=3.7pymc>=5.0 (Bayesian workflows)arviz>=0.16 (Bayesian diagnostics/plots)The example below is designed to be runnable end-to-end: it generates synthetic data, checks assumptions, runs an independent-samples t-test with effect size, performs power analysis, and prints an APA-style result string.
import numpy as np
import pandas as pd
import pingouin as pg
from statsmodels.stats.power import tt_ind_solve_power
# If your repo provides this module, use it; otherwise comment it out.
from scripts.assumption_checks import comprehensive_assumption_check
# ----------------------------
# 1) Create example dataset
# ----------------------------
rng = np.random.default_rng(7)
n_a, n_b = 50, 52
group_a = rng.normal(loc=75, scale=9, size=n_a)
group_b = rng.normal(loc=69, scale=9, size=n_b)
df = pd.DataFrame({
"score": np.r_[group_a, group_b],
"group": ["A"] * n_a + ["B"] * n_b
})
# ----------------------------
# 2) Assumption checks
# ----------------------------
assump = comprehensive_assumption_check(
data=df,
value_col="score",
group_col="group",
alpha=0.05
)
print("Assumption check summary:")
print(assump["summary"] if "summary" in assump else assump)
# ----------------------------
# 3) Run test + effect size
# ----------------------------
res = pg.ttest(group_a, group_b, correction="auto") # Welch if needed
t_stat = float(res["T"].iloc[0])
dfree = float(res["dof"].iloc[0])
pval = float(res["p-val"].iloc[0])
d = float(res["cohen-d"].iloc[0])
ci_low, ci_high = res["CI95%"].iloc[0]
# ----------------------------
# 4) Power analysis (planning)
# ----------------------------
n_required = tt_ind_solve_power(
effect_size=0.5, alpha=0.05, power=0.80, ratio=1.0, alternative="two-sided"
)
# ----------------------------
# 5) APA-style reporting string
# ----------------------------
m_a, sd_a = group_a.mean(), group_a.std(ddof=1)
m_b, sd_b = group_b.mean(), group_b.std(ddof=1)
apa = (
f"Group A (n = {n_a}, M = {m_a:.2f}, SD = {sd_a:.2f}) and "
f"Group B (n = {n_b}, M = {m_b:.2f}, SD = {sd_b:.2f}) differed, "
f"t({dfree:.0f}) = {t_stat:.2f}, p = {pval:.3f}, d = {d:.2f}, "
f"95% CI [{ci_low:.2f}, {ci_high:.2f}]."
)
print("\nAPA-style result:")
print(apa)
print(f"\nPlanning note: to detect d = 0.50 with 80% power, "
f"required n per group ≈ {n_required:.0f}.")Use references/test_selection_guide.md as the primary decision aid. The selection is typically driven by:
Common mappings:
The automated workflow in scripts/assumption_checks.py (referenced in the original documentation) is expected to cover:
Key parameter:
alpha (default commonly 0.05): decision threshold for assumption tests.Recommended remedies (see references/assumptions_and_diagnostics.md):
Effect sizes should be reported alongside inferential results (see references/effect_sizes_and_power.md):
Always prefer confidence intervals (frequentist) or credible intervals (Bayesian) to communicate precision.
Implemented via statsmodels.stats.power:
n given target effect size, alpha, and desired power.n, alpha, and desired power.Avoid “post-hoc power” computed from observed p-values; use sensitivity analysis instead.
Use references/reporting_standards.md to ensure inclusion of:
For Bayesian reporting (see references/bayesian_statistics.md), include:
| Field | Required | Format/Source | Example | If Missing |
|---|---|---|---|---|
| User task description | Yes | Text | Research question, writing goal, analysis objective | Stop and ask user to provide |
| Primary input material | Depends on task | Text, file path, ID, table, or literature | PMID, PDF, CSV, DOCX, keywords, etc. | Specify which material type is missing |
| Output preference | No | Text | Language, format, target journal, template | Use skill default format |
This skill accepts requests that match the documented purpose of statistical-analysis and include enough context to complete the workflow safely.
Do not continue the workflow when the request is out of scope, missing a critical input, or would require unsupported assumptions. Instead respond:
statistical-analysisonly handles its documented workflow. Please provide the missing required inputs or switch to a more suitable skill.
f5ef65b
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.