CtrlK
BlogDocsLog inGet started
Tessl Logo

senior-ml-engineer

ML engineering skill for productionizing models, building MLOps pipelines, and integrating LLMs. Covers model deployment, feature stores, drift monitoring, RAG systems, and cost optimization. Use when the user asks about deploying ML models to production, setting up MLOps infrastructure (MLflow, Kubeflow, Kubernetes, Docker), monitoring model performance or drift, building RAG pipelines, or integrating LLM APIs with retry logic and cost controls. Focused on production and operational concerns rather than model research or initial training.

69

Quality

85%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

The risk profile of this skill

SKILL.md
Quality
Evals
Security

Quality

Content

78%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A tightly structured, actionable body that pairs numbered workflows with executable code, tables, and validation checkpoints, supported by a clean one-level reference bundle. Recovery/rollback guidance and a few incomplete code imports are the main gaps.

Suggestions

Add an explicit 'if validation fails: rollback / halt promotion' step to the deployment and retraining workflows to complete the feedback loop.

Complete the Feast FeatureView example with the missing imports (e.g. ValueType, timedelta) so the snippet is copy-paste runnable.

Tighten the cost-management prose and the per-workflow introductory lines to trim tokens where the numbered steps already convey the intent.

DimensionReasoningScore

Conciseness

Dense and largely lean — numbered steps, tables, and code dominate rather than prose — but a few sections (e.g. the cost-management paragraph and per-workflow intro lines) could be tightened slightly.

4 / 5

Actionability

Provides executable Dockerfiles, Python snippets, and concrete script invocations ('python scripts/model_deployment_pipeline.py --model model.pkl --target staging'), though a couple of code samples (e.g. the Feast FeatureView) omit imports/context needed to run as-is.

4 / 5

Workflow Clarity

Each workflow is an 8-step sequence with an explicit bolded **Validation:** checkpoint, and the deployment flow includes a canary-then-promote feedback loop; minor gap is the absence of explicit rollback/recovery steps when validation fails.

4 / 5

Progressive Disclosure

Clear overview with a table of contents, well-signaled one-level-deep references (three reference files and three scripts, each summarized inline), and content appropriately split across the bundle — easy to navigate.

5 / 5

Total

17

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A well-crafted description with concrete capabilities, explicit use-when triggers, and a helpful scope boundary that distinguishes it from adjacent research skills. Minor synonym coverage in trigger terms is the only gap.

DimensionReasoningScore

Specificity

Lists multiple concrete actions — 'productionizing models, building MLOps pipelines, and integrating LLMs' plus 'model deployment, feature stores, drift monitoring, RAG systems, and cost optimization' — giving comprehensive coverage of the domain.

5 / 5

Completeness

Explicitly answers both 'what' (productionizing models, MLOps pipelines, LLM integration) and 'when' via a concrete 'Use when the user asks about...' clause with multiple trigger scenarios.

5 / 5

Trigger Term Quality

Strong natural keyword coverage including tool names ('MLflow, Kubeflow, Kubernetes, Docker') and phrases like 'deploying ML models to production' and 'building RAG pipelines', but a few common synonyms (e.g. 'inference', 'serving') are absent.

4 / 5

Distinctiveness Conflict Risk

Clear production-ops niche with an explicit boundary clause ('Focused on production and operational concerns rather than model research or initial training') minimizing overlap with research/training skills.

5 / 5

Total

19

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
alirezarezvani/claude-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.