CtrlK
BlogDocsLog inGet started
Tessl Logo

senior-data-scientist

World-class senior data scientist skill specialising in statistical modeling, experiment design, causal inference, and predictive analytics. Covers A/B testing (sample sizing, two-proportion z-tests, Bonferroni correction), difference-in-differences, feature engineering pipelines (Scikit-learn, XGBoost), cross-validated model evaluation (AUC-ROC, AUC-PR, SHAP), and MLflow experiment tracking — using Python (NumPy, Pandas, Scikit-learn), R, and SQL. Use when designing or analysing controlled experiments, building and evaluating classification or regression models, performing causal analysis on observational data, engineering features for structured tabular datasets, or translating statistical findings into data-driven business decisions.

72

Quality

88%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

The canonical home for this skill is senior-data-scientist in alirezarezvani/claude-skills

SKILL.md
Quality
Evals
Security

Quality

Content

85%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

Highly actionable, executable content with clear validated workflows, but progressive disclosure is weak: bulk implementations are inlined in SKILL.md and all referenced reference/script files are missing from the bundle.

Suggestions

Create the referenced bundle files (references/statistical_methods_advanced.md, experiment_design_frameworks.md, feature_engineering_patterns.md and scripts/*.py) or remove the broken references.

Move the long inline implementations into the reference files, keeping only concise quick-start snippets in SKILL.md so it serves as an overview.

Trim docstring lines that restate parameter values already obvious to a competent data scientist to tighten token usage.

DimensionReasoningScore

Conciseness

Mostly efficient with lean code and tight checklists, but docstrings restate obvious values ("baseline_rate: current conversion rate (e.g. 0.10)") and four long inline implementations could be trimmed or moved to references; minor over-explanation keeps it below a 5.

4 / 5

Actionability

Fully executable, copy-paste-ready Python covering the common cases (sample sizing, two-proportion z-test, ColumnTransformer pipeline, cross-validated evaluation with overfit-gap, OLS DiD with HC3), each with runnable functions and concrete checklists.

5 / 5

Workflow Clarity

Each workflow has a numbered checklist with explicit validation checkpoints (SRM check, overfit_gap > 0.05, Bonferroni, parallel-trends validation), and the batch pipeline commands include a status-based fail-fast feedback loop ("any status other than 'completed' means the stage failed — fix before moving to the next pipeline stage").

5 / 5

Progressive Disclosure

Sections and reference links are present, but the body inlines four full code implementations that arguably belong in separate files, and every referenced path (references/*.md, scripts/*.py) points to files that do not exist in the bundle — references are signaled but not actually provided.

3 / 5

Total

17

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, specific description that clearly states both capabilities and trigger conditions in third person, with comprehensive concrete actions. Trigger terms could be slightly less jargon-heavy and include more natural synonyms.

DimensionReasoningScore

Specificity

Lists multiple concrete actions ("sample sizing, two-proportion z-tests, Bonferroni correction", "difference-in-differences", "cross-validated model evaluation (AUC-ROC, AUC-PR, SHAP)", "MLflow experiment tracking") with comprehensive coverage across experiment design, causal inference, and ML.

5 / 5

Completeness

Explicitly answers both what ("Covers A/B testing ... MLflow experiment tracking") and when ("Use when designing or analysing controlled experiments, building and evaluating ... models") with concrete trigger phrases.

5 / 5

Trigger Term Quality

Good natural coverage ("A/B testing", "classification or regression models", "feature engineering", "causal analysis") but leans technical ("two-proportion z-tests", "Bonferroni correction") and omits common synonyms a non-specialist user might say; not quite comprehensive enough for a 5.

4 / 5

Distinctiveness Conflict Risk

Occupies a clear senior-data-scientist niche with distinct, specific triggers (experiment design, causal inference, model evaluation) that minimize overlap with general coding or data skills.

5 / 5

Total

19

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

referenced_paths_exist

Referenced path issues: 6 missing

Warning

Total

15

/

16

Passed

Repository
alirezarezvani/claude-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.