CtrlK
BlogDocsLog inGet started
Tessl Logo

ml-pipeline-workflow

Build end-to-end MLOps pipelines from data preparation through model training, validation, and production deployment. Use when creating ML pipelines, implementing MLOps practices, or automating model training and deployment workflows.

69

0.98x
Quality

61%

Does it follow best practices?

Impact

73%

0.98x

Average score across 3 eval scenarios

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./tests/ext_conformance/artifacts/agents-wshobson/machine-learning-ops/skills/ml-pipeline-workflow/SKILL.md

The canonical home for this skill is ml-pipeline-workflow in wshobson/agents

SKILL.md
Quality
Evals
Security

Quality

Content

38%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is well-structured and describes a coherent pipeline workflow, but it under-delivers on execution: code examples are comment placeholders, all referenced bundle files are missing, and much of the content restates generic MLOps knowledge. Validation appears as a phase but without explicit fail-and-retry feedback loops.

Suggestions

Create the referenced files (references/data-preparation.md, model-training.md, model-validation.md, model-deployment.md, and the three assets) or remove the references — currently every deep-dive pointer in the body is broken.

Replace placeholder code blocks ("# See assets/pipeline-dag.yaml.template for full example") with executable snippets, e.g. a minimal Airflow DAG or a working pipeline-stage configuration.

Trim generic knowledge sections (Integration Points, Best Practices, Deployment Strategies) to only non-obvious guidance, and add explicit validate-and-rollback feedback loops to the Production Workflow for the deployment phase.

DimensionReasoningScore

Conciseness

Large portions restate MLOps knowledge Claude already has — e.g. "Start with shadow deployments", "Use canary releases for validation", "Use model registries (MLflow, Weights & Biases)", and tool catalogs for Airflow/Dagster/Kubeflow/SageMaker/Vertex AI. Several padded sections (Integration Points, Best Practices) add little beyond common knowledge; not a 1 because there is no tutorial-style pedagogy and section organization is clean.

2 / 5

Actionability

Concrete elements exist (the six-stage list, the YAML stage/dependencies snippet, the four-phase workflow) but the code blocks are non-executable placeholders — "# 2. Configure dependencies / # See assets/pipeline-dag.yaml.template for full example" and "# Stream processing for real-time features" — and the referenced asset files do not exist. This matches the 3 anchor (guidance present, pseudocode/placeholder instead of executable code, key details missing); not 2 because stage definitions and the YAML example are genuinely concrete.

3 / 5

Workflow Clarity

The Production Workflow gives a clear four-phase sequence and includes a Validation phase with "Run validation test suite" and "Approve for deployment", but there are no explicit feedback loops (no "if validation fails, rollback/fix and retry" step) for deployment, a risky batch operation. Per the rubric's feedback-loop guidance this caps the score at 3; it is above 2 because the sequence itself is well defined with a dedicated validation phase.

3 / 5

Progressive Disclosure

The body repeatedly points to "references/data-preparation.md", "references/model-training.md", "assets/pipeline-dag.yaml.template", and "assets/training-config.yaml", but no references/, scripts/, or assets/ directories exist — every deep-dive pointer is broken. Meanwhile content that belongs in those files (best-practice lists, tool catalogs) is inlined. This matches the 2 anchor (content that belongs in separate files is inlined, references unusable); not 1 because section headers and one-level reference signaling are present.

2 / 5

Total

10

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description that explicitly covers both what the skill does and when to use it, with natural trigger phrases in third person. Minor room to improve by adding common synonym terms and sharpening the action list.

DimensionReasoningScore

Specificity

"Build end-to-end MLOps pipelines from data preparation through model training, validation, and production deployment" names the domain plus several concrete pipeline-stage actions (preparation, training, validation, deployment). It falls short of the 5 anchor because the actions are stage names rather than the multiple distinct concrete operations in the anchor example, and exceeds 3 because coverage spans more than 1-2 actions.

4 / 5

Completeness

Both parts are explicit: the "what" is "Build end-to-end MLOps pipelines from data preparation through model training, validation, and production deployment" and the "when" is an explicit "Use when..." clause with three concrete trigger phrases. This matches the 5 anchor's pattern of clearly answering both with concrete triggers, in third person.

5 / 5

Trigger Term Quality

"Use when creating ML pipelines, implementing MLOps practices, or automating model training and deployment workflows" supplies several natural phrases a user would say. Not 5 because common variations like "machine learning pipeline", "CI/CD for models", or "model serving" are missing; not 3 because multiple natural terms are present, not just one generic keyword.

4 / 5

Distinctiveness Conflict Risk

The MLOps pipeline niche is mostly distinct with clear triggers, but "automating model training and deployment workflows" overlaps with closely related skills the body itself lists (model-deployment-patterns, experiment-tracking-setup). Not 5 due to that minor overlap risk; clearly above 3 since the scope is a recognizable niche rather than generic.

4 / 5

Total

17

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

referenced_paths_exist

Referenced path issues: 4 missing

Warning

Total

15

/

16

Passed

Repository
Dicklesworthstone/pi_agent_rust
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.