CtrlK
BlogDocsLog inGet started
Tessl Logo

senior-ml-engineer

World-class ML engineering skill for productionizing ML models, MLOps, and building scalable ML systems. Expertise in PyTorch, TensorFlow, model deployment, feature stores, model monitoring, and ML infrastructure. Includes LLM integration, fine-tuning, RAG systems, and agentic AI. Use when deploying ML models, building ML platforms, implementing MLOps, or integrating LLMs into production systems.

60

Quality

70%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./bundled/skills/senior-ml-engineer/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

57%Weight 40%Scale 1-3

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The skill is well-structured for progressive disclosure, with real one-level-deep references and executable commands. Its weaknesses are verbosity from generic senior-engineer platitudes and a lack of validated, sequenced workflows for the risky deployment operations it describes.

Suggestions

Remove the generic filler sections ('Core Expertise', 'Senior-Level Responsibilities', 'Best Practices') that restate what Claude already knows, to improve token efficiency.

Turn the deployment commands into an explicit sequenced workflow with a validation checkpoint (e.g., validate manifests / run health_check.py) before declaring success.

Replace abstract capability bullets ('A/B testing infrastructure', 'Feature store integration') with a concrete executable example or a pointer into the reference files.

DimensionReasoningScore

Conciseness

Content is bulleted rather than prose-explaining basics, but large sections ('Core Expertise', 'Best Practices', 'Senior-Level Responsibilities') are generic platitudes Claude already knows ('Documentation as code', 'Mentor junior engineers'), fitting 'mostly efficient but includes unnecessary explanation'; not 3 because many tokens do not earn their place, not 1 because it avoids explaining basic concepts in prose.

2 / 3

Actionability

Provides real executable commands ('python scripts/model_deployment_pipeline.py --input data/ --output results/', 'docker build -t service:v1 .', 'kubectl apply -f k8s/') but they are outnumbered by abstract capability descriptions ('Model serving with low latency', 'A/B testing infrastructure'), matching the level-2 'some concrete guidance but incomplete' anchor; not 3 because most content describes rather than instructs, not 1 because the CLI/script commands are genuinely executable and reference real bundle files.

2 / 3

Workflow Clarity

Commands are grouped (Development, Training, Deployment, Monitoring) giving a loose sequence, but there are no validation checkpoints or error-recovery feedback loops for the risky deployment operations, which caps the score at 2 per the rubric's destructive/batch guideline; not 3 because validation steps are absent, not 1 because some structure is present rather than missing or unclear steps.

2 / 3

Progressive Disclosure

The body is an overview pointing to real one-level-deep references ('references/mlops_production_patterns.md', 'references/llm_integration_guide.md', 'references/rag_system_architecture.md') and a 'scripts/' directory, all verified to exist and clearly signaled in both the 'Reference Documentation' and 'Resources' sections, matching the level-3 anchor; not 2 because the references are real, one level deep, and well-navigated rather than monolithic or nested.

3 / 3

Total

9

/

12

Passed

Description

82%Weight 40%Scale 1-3

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is strong on completeness and trigger-term quality, with an explicit 'Use when' clause and natural keywords. It is somewhat weakened by buzzword tech-stack padding and a broad scope that raises overlap risk with adjacent skills.

Suggestions

Trim the framework name-lists ('PyTorch, TensorFlow, Scikit-learn, XGBoost...') in favor of a few concrete representative actions to tighten specificity.

Narrow the scope or lead with a single primary trigger so the skill is less likely to conflict with dedicated LLM or MLOps skills.

DimensionReasoningScore

Specificity

States concrete actions ('productionizing ML models', 'building scalable ML systems', 'deploying', 'integrating LLMs') but pads them with tech-stack name-drops ('PyTorch, TensorFlow... feature stores') rather than crisp concrete actions, fitting 'names domain and some actions' rather than the level-3 list of specific actions.

2 / 3

Completeness

Explicitly answers 'what' (productionizing, MLOps, LLM integration, RAG, fine-tuning) and 'when' via an explicit 'Use when...' clause, matching the level-3 anchor for both what and when.

3 / 3

Trigger Term Quality

'Use when deploying ML models, building ML platforms, implementing MLOps, or integrating LLMs' are natural phrases a user would say, giving good coverage of common trigger terms; not below 3 because the variations are well covered.

3 / 3

Distinctiveness Conflict Risk

The breadth (ML engineering + MLOps + LLMs + RAG + agentic AI) could overlap with dedicated LLM, data-engineering, or MLOps skills; somewhat specific but not a clear single niche, so it sits at level 2 rather than 3.

2 / 3

Total

10

/

12

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

referenced_paths_exist

Referenced path issues: 3 missing

Warning

Total

15

/

16

Passed

Repository
foryourhealth111-pixel/Vibe-Skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.