CtrlK
BlogDocsLog inGet started
Tessl Logo

senior-ml-engineer

ML engineering skill for productionizing models, building MLOps pipelines, and integrating LLMs. Covers model deployment, feature stores, drift monitoring, RAG systems, and cost optimization. Use when the user asks about deploying ML models to production, setting up MLOps infrastructure (MLflow, Kubeflow, Kubernetes, Docker), monitoring model performance or drift, building RAG pipelines, or integrating LLM APIs with retry logic and cost controls. Focused on production and operational concerns rather than model research or initial training.

67

Quality

84%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

The canonical home for this skill is senior-ml-engineer in alirezarezvani/claude-skills

SKILL.md
Quality
Evals
Security

Quality

Content

71%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a well-structured, action-dense overview: five numbered workflows each with explicit validation criteria, decision tables with concrete thresholds, and executable code templates. Its main weaknesses are minor trimmable padding, a code snippet with missing imports, and — most significantly — a progressive-disclosure layer whose six referenced bundle files do not actually exist in the skill directory.

Suggestions

Ship the referenced bundle files (references/mlops_production_patterns.md, references/llm_integration_guide.md, references/rag_system_architecture.md, and the three scripts/*.py) or remove those sections — as written, every referenced path is dangling and the detailed content is unreachable.

Trim the Tech Stack table and reduce the serving/vector-database comparison tables to the decision-relevant columns, since Claude already knows these tools; keep only the curated guidance (latency/throughput picks, thresholds) that it does not.

Fix the Feast FeatureView snippet to be executable as written (add the ValueType and timedelta imports), and replace the cross-skill path to engineering-team/skills/senior-prompt-engineer/scripts/prompt_optimizer.py with a self-contained example, since a path into another skill is fragile.

DimensionReasoningScore

Conciseness

The body is dense and table/code-driven with almost no prose padding, and thresholds like "p95 latency < 100ms, error rate < 0.1%" earn their tokens. Not 5 because the Tech Stack table ("PyTorch, TensorFlow, Scikit-learn, XGBoost…") and parts of the serving/vector-DB comparison tables restate common knowledge Claude already has and could be trimmed; not 3 because there are no concept explanations, only curated decision matrices.

4 / 5

Actionability

Mostly executable guidance: a complete Dockerfile template, runnable drift-detection code ("statistic, p_value = ks_2samp(reference, current)"), and concrete script invocations ("python scripts/model_deployment_pipeline.py --model model.pkl --target staging"). Not 5 because the Feast snippet is missing imports (ValueType, timedelta) and the referenced scripts cannot be verified; not 3 because the provided code and commands are concrete, not pseudocode.

4 / 5

Workflow Clarity

All five workflows are clearly sequenced as numbered 8-step lists, each ending with a bolded validation step carrying hard criteria ("p95 latency < 100ms", "PSI > 0.2", "cost within budget"), and the retraining-trigger table maps detection to action. Not 5 because error-recovery feedback loops are only implied ("Promote to full production if metrics pass") — no step says what to do when validation fails; not 3 because explicit validation checkpoints are present in every workflow.

4 / 5

Progressive Disclosure

The in-file structure is good (table of contents, well-signaled reference sections with content summaries like "references/llm_integration_guide.md contains: Provider abstraction layer pattern…"), but scoring against the actual bundle reveals no references/, scripts/, or assets/ directories exist — all six referenced paths are dangling, so the detailed material is unreachable. Not 4 because the bundle does not back the clearly-signaled references; not 2 because the SKILL.md itself is well-organized with references properly surfaced rather than buried or inlined.

3 / 5

Total

15

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: third-person voice, concrete actions, comprehensive domain coverage, an explicit 'Use when' clause with specific trigger phrases, and a clear boundary statement excluding model research and initial training. The only weakness is that a few of the skill's own workflow terms (model serving, automated retraining, A/B testing) are absent from the description field itself.

DimensionReasoningScore

Specificity

Lists multiple concrete actions ("productionizing models, building MLOps pipelines, and integrating LLMs") plus comprehensive coverage of the domain ("model deployment, feature stores, drift monitoring, RAG systems, and cost optimization"). Not 4 because there is no material gap in coverage — the listed actions span every workflow the skill body implements.

5 / 5

Completeness

Explicitly answers both questions: what ("ML engineering skill for productionizing models, building MLOps pipelines, and integrating LLMs") and when ("Use when the user asks about deploying ML models to production…") with concrete trigger phrases. Not 4 because the 'when' clause is already fully explicit and specific — nothing could be made clearer.

5 / 5

Trigger Term Quality

Strong natural phrases users would say: "deploying ML models to production", "setting up MLOps infrastructure (MLflow, Kubeflow, Kubernetes, Docker)", "monitoring model performance or drift", "building RAG pipelines", "integrating LLM APIs". Not 5 because a few natural terms are missing from the description itself — "model serving", "automated retraining", and "A/B testing" are skill workflows that only appear in the triggers list, not the description field.

4 / 5

Distinctiveness Conflict Risk

Clear niche with distinct triggers (MLOps, drift, RAG pipelines, LLM integration) and an explicit boundary sentence ("Focused on production and operational concerns rather than model research or initial training") that prevents overlap with ML research or general DevOps skills. Not 4 because the boundary is stated explicitly, leaving only negligible overlap risk.

5 / 5

Total

19

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

referenced_paths_exist

Referenced path issues: 6 missing

Warning

Total

14

/

16

Passed

Repository
alirezarezvani/claude-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.