CtrlK
BlogDocsLog inGet started
Tessl Logo

senior-data-scientist

World-class senior data scientist skill specialising in statistical modeling, experiment design, causal inference, and predictive analytics. Covers A/B testing (sample sizing, two-proportion z-tests, Bonferroni correction), difference-in-differences, feature engineering pipelines (Scikit-learn, XGBoost), cross-validated model evaluation (AUC-ROC, AUC-PR, SHAP), and MLflow experiment tracking — using Python (NumPy, Pandas, Scikit-learn), R, and SQL. Use when designing or analysing controlled experiments, building and evaluating classification or regression models, performing causal analysis on observational data, engineering features for structured tabular datasets, or translating statistical findings into data-driven business decisions.

70

Quality

88%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

The canonical home for this skill is senior-data-scientist in alirezarezvani/claude-skills

SKILL.md
Quality
Evals
Security

Quality

Content

85%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is highly actionable — four complete, executable workflow recipes each capped with a validation checklist and an explicit failure-feedback loop for the bundled pipelines. Its one real defect is progressive disclosure: every referenced bundle path (three references/*.md files and three scripts/*.py files) is a dead link because those files were never shipped alongside SKILL.md.

Suggestions

Ship the referenced bundle files — create references/statistical_methods_advanced.md, references/experiment_design_frameworks.md, references/feature_engineering_patterns.md and scripts/experiment_designer.py, scripts/feature_engineering_pipeline.py, scripts/model_evaluation_suite.py — or delete the "Reference Documentation" section and the script-invocation lines; as written, all six referenced paths 404.

Slim SKILL.md to a concise overview (one short signature example per workflow) and move the full function bodies into the reference files, so the ~200-line monolithic body follows the progressive-disclosure structure it already advertises.

Replace the buzzword opener ("World-class senior data scientist skill for production-grade AI/ML/Data systems.") and the generic pytest/black/pylint boilerplate with a one-line statement of what the skill's own workflows do.

DimensionReasoningScore

Conciseness

The body is dominated by dense, executable code with directive comments rather than explanations of concepts Claude already knows, e.g. the checklists "Never fit transformers on the full dataset — fit on train, transform test". Not 5 because of minor trimmable material: the filler opener "World-class senior data scientist skill for production-grade AI/ML/Data systems." and generic boilerplate commands ("python -m pytest tests/ -v --cov=src/", "python -m black src/ && python -m pylint src/"); not 3 because there is no substantive padding beyond these.

4 / 5

Actionability

All four workflows are complete, copy-paste-ready functions with imports, parameters, and return values (e.g. calculate_sample_size, analyze_experiment, build_feature_pipeline, diff_in_diff), plus concrete CLI invocations with explicit input/output arguments ("--input experiment_spec.json --output experiment_design.json"). This matches the anchor: fully executable code covering the common cases; not 4 because there are no gaps in the code itself.

5 / 5

Workflow Clarity

Each workflow ends with a numbered checklist containing explicit validation checkpoints ("Check for sample ratio mismatch: abs(n_control - n_treatment) / expected < 0.01", "Check overfit_gap > 0.05", "Run a baseline (e.g. DummyClassifier) and verify the model beats it", "Validate parallel trends in pre-period"), and the pipeline section defines an explicit failure feedback loop: "any status other than 'completed' means the stage failed — fix before moving to the next pipeline stage". Not 4 because both checklists for complex processes and error-recovery loops are present, matching the top anchor.

5 / 5

Progressive Disclosure

The "Reference Documentation" section clearly signals three reference files and "Common Commands" invokes three bundled scripts, but none of these exist — there is no references/ or scripts/ directory in the bundle, so all six referenced paths are dead links. Not 4 because missing files entirely is more than "minor organization gaps": navigation is broken and the ~200-line body carries all substantive content monolithically; not 2 because section structure and reference signaling are good, only the targets are absent.

3 / 5

Total

17

/

20

Passed

Description

88%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: explicit and comprehensive "what" plus a concrete "Use when..." trigger clause in third-person voice, with only minor gaps in keyword synonyms. The lone weakness is the buzzword opener "World-class senior data scientist skill specialising in..." which adds fluff without adding trigger value.

DimensionReasoningScore

Specificity

The description lists multiple concrete actions with sub-specifics: "Covers A/B testing (sample sizing, two-proportion z-tests, Bonferroni correction), difference-in-differences, feature engineering pipelines (Scikit-learn, XGBoost), cross-validated model evaluation (AUC-ROC, AUC-PR, SHAP), and MLflow experiment tracking" — five action areas, each named with concrete techniques. Not 4 because coverage is comprehensive rather than having minor gaps; the one buzzword opener ("World-class") does not dilute the concrete action list.

5 / 5

Completeness

Both questions are explicitly answered: the "what" via "Covers A/B testing ... difference-in-differences ... MLflow experiment tracking", and the "when" via an explicit trigger clause — "Use when designing or analysing controlled experiments, building and evaluating classification or regression models, performing causal analysis on observational data, engineering features for structured tabular datasets, or translating statistical findings into data-driven business decisions". This matches the anchor-5 example pattern exactly; not 4 because the "when" clause is already explicit and concrete.

5 / 5

Trigger Term Quality

Natural phrases users would say are present: "A/B testing", "experiment design", "classification or regression models", "causal analysis", "structured tabular datasets", "data-driven business decisions". Not 5 because common variations are missing — e.g. "machine learning" / "ML" never appears, nor do "forecasting", "uplift", or tool-name triggers a user might naturally mention; not 3 because keyword coverage is genuinely good with several distinct natural trigger phrases.

4 / 5

Distinctiveness Conflict Risk

The niche is mostly distinct — "two-proportion z-tests, Bonferroni correction", "difference-in-differences", and "causal analysis on observational data" are specialized triggers unlikely to fire for unrelated skills. Not 5 because the field is broad: generic requests like "build a predictive model" or "analyze this dataset" could overlap with general ML-engineering or data-analysis skills; not 3 because the statistical/experimental-design focus keeps overlap minor rather than routine.

4 / 5

Total

18

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

referenced_paths_exist

Referenced path issues: 6 missing

Warning

Total

15

/

16

Passed

Repository
alirezarezvani/claude-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.