CtrlK
BlogDocsLog inGet started
Tessl Logo

senior-ml-engineer

World-class ML engineering skill for productionizing ML models, MLOps, and building scalable ML systems. Expertise in PyTorch, TensorFlow, model deployment, feature stores, model monitoring, and ML infrastructure. Includes LLM integration, fine-tuning, RAG systems, and agentic AI. Use when deploying ML models, building ML platforms, implementing MLOps, or integrating LLMs into production systems.

54

Quality

68%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./bundled/skills/senior-ml-engineer/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

40%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is well-structured with real references and some executable commands, but it is heavily padded with generic capability lists Claude already knows and lacks any sequenced workflow with validation checkpoints. It reads more like a topical outline than actionable guidance.

Suggestions

Cut the generic 'Core Expertise', 'Best Practices', and 'Senior-Level Responsibilities' bullet lists, which restate knowledge Claude already has, and keep only non-obvious specifics.

Add a sequenced end-to-end workflow (e.g., train -> validate -> package -> deploy -> monitor) with explicit validation checkpoints for the batch/destructive deployment steps.

Replace abstract pattern descriptions with concrete, copy-paste-ready examples (e.g., a real k8s manifest snippet or a working monitoring config) or move detail into the reference files.

DimensionReasoningScore

Conciseness

Large swaths of the body are generic padded bullet lists Claude already knows ('Test-driven development', 'Drive architectural decisions', 'Advanced production patterns and architectures'), adding little actionable knowledge beyond enumeration.

2 / 5

Actionability

Quick Start and Common Commands give genuinely executable commands (python scripts/..., docker build, kubectl apply, helm upgrade), but the majority of sections are abstract descriptions ('Model serving with low latency', 'Batching and caching strategies') with no concrete steps or code.

3 / 5

Workflow Clarity

No real multi-step workflow is sequenced; Quick Start is three independent commands and Common Commands is loosely grouped. Batch/destructive operations (kubectl apply, helm upgrade) appear with no validation or verification checkpoints.

2 / 5

Progressive Disclosure

Clear section structure with well-signaled, one-level-deep references to three real reference files and the scripts/ directory (all verified to exist). Minor gaps: the body inlines generic content that belongs in references, and the reference files themselves are thin.

4 / 5

Total

11

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is strong: it explicitly covers both what the skill does and when to use it with concrete trigger phrases, and names many specific capabilities. It loses points on specificity and distinctiveness due to some jargon/domain-noun phrasing and a broad scope that risks overlap with narrower skills.

Suggestions

Replace jargon/domain nouns ('productionizing', 'MLOps', 'ML infrastructure') with concrete actions (e.g., 'deploy models to Kubernetes', 'build retraining pipelines', 'set up drift monitoring').

Add common trigger synonyms users would naturally say, such as 'training models', 'serving inference', or 'model registry'.

Tighten the scope or add a distinguishing qualifier to reduce overlap risk with dedicated RAG or deployment skills.

DimensionReasoningScore

Specificity

Lists several specific capabilities (model deployment, feature stores, model monitoring, fine-tuning, RAG systems) but several items like 'productionizing', 'MLOps', and 'ML infrastructure' are domain nouns/jargon rather than crisp concrete actions, leaving minor gaps.

4 / 5

Completeness

It clearly states what the skill does (productionizing ML models, MLOps, LLM integration, fine-tuning, RAG) and gives an explicit 'Use when...' clause with multiple concrete trigger phrases, matching the anchor for clearly answering both what and when.

5 / 5

Trigger Term Quality

The 'Use when deploying ML models, building ML platforms, implementing MLOps, or integrating LLMs' clause provides good natural trigger phrases, but common variations like 'training models' or 'inference' are missing.

4 / 5

Distinctiveness Conflict Risk

The ML/AI engineering production niche has distinct triggers, but the very broad surface (MLOps, RAG, LLMs, deployment, monitoring) creates minor overlap risk with more specialized skills.

4 / 5

Total

17

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

referenced_paths_exist

Referenced path issues: 3 missing

Warning

Total

15

/

16

Passed

Repository
foryourhealth111-pixel/Vibe-Skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.