CtrlK
BlogDocsLog inGet started
Tessl Logo

senior-data-scientist

World-class data science skill for statistical modeling, experimentation, causal inference, and advanced analytics. Expertise in Python (NumPy, Pandas, Scikit-learn), R, SQL, statistical methods, A/B testing, time series, and business intelligence. Includes experiment design, feature engineering, model evaluation, and stakeholder communication. Use when designing experiments, building predictive models, performing causal analysis, or driving data-driven decisions.

67

1.41x
Quality

60%

Does it follow best practices?

Impact

68%

1.41x

Average score across 3 eval scenarios

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./bundled/skills/senior-data-scientist/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

40%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is structurally okay — real, clearly signaled reference files and a scripts directory — but it is padded with generic senior-engineer boilerplate, its Quick Start commands point to stub scripts, and it lacks any sequenced workflow with validation checkpoints for its destructive/batch operations.

Suggestions

Cut the generic boilerplate sections ('Core Expertise', 'Best Practices', 'Senior-Level Responsibilities') and keep only data-science-specific guidance Claude would not already know.

Replace the stub scripts (bodies are '# Implementation here') with real executable logic, or remove the commands until the tools actually work.

Add a sequenced end-to-end workflow (e.g., design experiment -> validate config -> run -> evaluate -> check metrics) with explicit validation checkpoints before any deploy/training step.

DimensionReasoningScore

Conciseness

Large sections ('Core Expertise', 'Best Practices', 'Senior-Level Responsibilities', 'Security & Compliance') are generic senior-engineer boilerplate Claude already knows ('Test-driven development', 'Monitor everything critical', 'Mentor junior engineers'), adding noticeable padding without data-science-specific signal.

2 / 5

Actionability

Concrete command invocations are present ('python scripts/experiment_designer.py --input data/ --output results/'), but they point at stub scripts whose bodies are '# Implementation here' / '# Add validation logic', and the rest are generic tool commands, leaving the guidance incomplete rather than copy-paste ready.

3 / 5

Workflow Clarity

There is no real sequenced workflow — 'Quick Start' lists three independent commands in parallel — and destructive/batch operations like '--deploy' and training carry no validation or verification checkpoints, so the content stays well below the cap of 3.

2 / 5

Progressive Disclosure

References to references/statistical_methods_advanced.md, experiment_design_frameworks.md, and feature_engineering_patterns.md are real, one level deep, and clearly signaled in dedicated 'Reference Documentation' and 'Resources' sections, though the body still carries generic inline content that could be trimmed.

4 / 5

Total

11

/

20

Passed

Description

80%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is strong: it states concrete capabilities and pairs an explicit 'what' with a concrete 'Use when...' trigger clause. Its main weakness is breadth and buzzword padding ('world-class'), which raises overlap risk and keeps specificity and distinctiveness below the top level.

DimensionReasoningScore

Specificity

Lists several concrete capabilities ('statistical modeling, experimentation, causal inference...A/B testing, time series...feature engineering, model evaluation') rather than vague language, though 'world-class' is fluff and the items are domain categories more than discrete operations, so it stops short of 5.

4 / 5

Completeness

It explicitly answers both 'what' (the enumerated expertise and capabilities) and 'when' (the 'Use when...' clause with concrete trigger phrases), matching the top anchor.

5 / 5

Trigger Term Quality

The 'Use when designing experiments, building predictive models, performing causal analysis, or driving data-driven decisions' clause supplies natural phrases a user would say, with good but not exhaustive synonym/extension coverage.

4 / 5

Distinctiveness Conflict Risk

The scope is a distinct but very broad data-science niche spanning modeling, experimentation, causal inference, BI, and MLOps, creating real overlap risk with general ML-engineering or analytics skills rather than a clean unique trigger surface.

3 / 5

Total

16

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

referenced_paths_exist

Referenced path issues: 3 missing

Warning

Total

15

/

16

Passed

Repository
foryourhealth111-pixel/Vibe-Skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.