CtrlK
BlogDocsLog inGet started
Tessl Logo

mlops-engineer

Build comprehensive ML pipelines, experiment tracking, and model registries with MLflow, Kubeflow, and modern MLOps tools.

36

Quality

32%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./plugins/AI-Agents-Safe-Coding-Skills/skills/mlops-engineer/SKILL.md

The canonical home for this skill is mlops-engineer in administrakt0r/AI-Agents-Safe-Coding-Skills

SKILL.md
Quality
Evals
Security

Quality

Content

3%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

This skill is essentially a persona description or job posting rather than actionable instructions. It consists almost entirely of bullet-point lists of tools and abstract capabilities that Claude already knows, with zero executable code, no concrete workflows, and no validation steps. The content would need to be fundamentally restructured to provide actual value as a skill.

Suggestions

Replace the extensive tool/capability lists with concrete, executable code examples for the most common MLOps tasks (e.g., setting up an MLflow experiment, creating a Kubeflow pipeline YAML, writing a Dockerfile for model serving).

Add explicit multi-step workflows with validation checkpoints for critical operations like model deployment, pipeline creation, and infrastructure provisioning (e.g., 'terraform plan' before 'terraform apply').

Remove the 'Behavioral Traits', 'Knowledge Base', 'Capabilities' lists, and 'Example Interactions' sections — these describe what Claude already knows and waste token budget.

Move detailed reference material into the referenced `resources/implementation-playbook.md` bundle file and keep SKILL.md as a concise overview with only the most essential patterns and commands.

DimensionReasoningScore

Conciseness

Extremely verbose with extensive lists of tools, platforms, and capabilities that Claude already knows. The content reads like a resume or job description rather than actionable instructions. Massive sections like 'Capabilities', 'Knowledge Base', and 'Behavioral Traits' are pure padding that explain concepts Claude is already familiar with.

1 / 5

Actionability

No concrete code, commands, or executable examples anywhere in the skill. The entire content is abstract descriptions and bullet-point lists of tools and concepts. The 'Response Approach' section provides only vague high-level steps like 'Analyze MLOps requirements' with no specifics on how to actually do anything.

1 / 5

Workflow Clarity

No clear multi-step workflow with validation checkpoints. The 'Response Approach' lists 8 generic steps but they are abstract directives without concrete commands, validation steps, or error recovery. For a skill involving infrastructure provisioning and deployment (destructive/batch operations), the complete absence of validation steps is a critical gap.

1 / 5

Progressive Disclosure

References `resources/implementation-playbook.md` but no bundle files are provided, so the reference is unverifiable. The massive amount of content (tool lists, capabilities, behavioral traits) is all inlined in a monolithic fashion when it clearly should be split into separate reference files or removed entirely.

2 / 5

Total

5

/

20

Passed

Description

61%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description identifies a clear domain (MLOps) and names specific tools (MLflow, Kubeflow), which aids distinctiveness and trigger matching. However, it lacks a 'Use when...' clause, which is a significant gap for skill selection, and the capabilities listed are more categorical than action-specific. Adding explicit trigger guidance and more concrete actions would substantially improve it.

Suggestions

Add a 'Use when...' clause with trigger phrases like 'Use when the user asks about ML pipelines, experiment tracking, model versioning, MLflow setup, Kubeflow configuration, or MLOps workflows'.

Replace broad category terms with specific actions, e.g., 'Set up experiment tracking with MLflow, build training pipelines in Kubeflow, configure model registries, automate model deployment and monitoring'.

Include additional natural trigger terms and synonyms such as 'model deployment', 'model serving', 'ML CI/CD', 'training pipeline', and 'model versioning'.

DimensionReasoningScore

Specificity

Names the domain (ML pipelines/MLOps) and lists a few concrete concepts (experiment tracking, model registries), but these are more like categories than specific actions. It doesn't describe concrete actions like 'configure hyperparameter sweeps' or 'deploy models to production endpoints'.

3 / 5

Completeness

Has a clear 'what' (build ML pipelines, experiment tracking, model registries with specific tools), but completely lacks a 'when' clause. There is no 'Use when...' guidance, which per the rubric caps completeness at 3.

3 / 5

Trigger Term Quality

Includes several natural keywords users would say: 'ML pipelines', 'experiment tracking', 'model registries', 'MLflow', 'Kubeflow', 'MLOps'. Missing some common variations like 'model deployment', 'model serving', 'feature store', 'CI/CD for ML', or 'machine learning operations'.

4 / 5

Distinctiveness Conflict Risk

The mention of specific tools (MLflow, Kubeflow) and the MLOps focus makes it fairly distinct. However, it could overlap with general ML/data science skills or DevOps/infrastructure skills. The tool-specific mentions reduce conflict risk considerably.

4 / 5

Total

14

/

20

Passed

Validation

90%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation10 / 11 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

10

/

11

Passed

Repository
administrakt0r/AI-Agents-Safe-Coding-Skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.