CtrlK
BlogDocsLog inGet started
Tessl Logo

ml-pipeline-automation

Automate ML workflows with Airflow, Kubeflow, MLflow. Use for reproducible pipelines, retraining schedules, MLOps, or encountering task failures, dependency errors, experiment tracking issues.

83

1.25x
Quality

75%

Does it follow best practices?

Impact

100%

1.25x

Average score across 3 eval scenarios

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./plugins/ml-pipeline-automation/skills/ml-pipeline-automation/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

67%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A dense, highly actionable skill body with good structure and a well-signaled reference layer. Its main weaknesses are token cost — duplicated DAG examples, a tools table for tools the skill doesn't cover, and known-issues detail that belongs in references — plus a few placeholder code snippets and version-pinned install lines that will age.

Suggestions

Remove the duplicated train_model DAG (keep one full example; the Quick Start can reuse it instead of repeating it inline).

Move most of the seven 'Known Issues Prevention' sections with their full code into references/airflow-patterns.md, keeping one-line problem pointers in SKILL.md.

Drop the version-pinned install commands and December-2025 date notes in favor of an unpinned install plus a 'check latest on PyPI' note, and complete placeholder snippets (train_rf/task_group example, my_function) or mark them explicitly as patterns.

DimensionReasoningScore

Conciseness

The ~430-line body is mostly efficient code with little prose padding, but it has real redundancy: the train_model MLflow DAG appears twice (Quick Start and Basic Airflow DAG), the tools table covers Prefect/Dagster which the skill does not otherwise address, and seven full 'Known Issues Prevention' sections with complete code examples belong in references. Time-sensitive version pins ('pip install apache-airflow==3.1.5... current as of December 2025') further cost tokens and will age. Not a 2 because there is no concept-explaining filler prose like the bad examples; not a 4 because the duplication and inline bulk are clearly trimmable.

3 / 5

Actionability

Nearly all guidance is copy-paste executable: a 5-step quick start with a complete runnable DAG, XCom validation code, timeout/alert configurations, and sensor patterns. Minor gaps keep it from 5: some snippets are placeholders ('train_rf = PythonOperator(task_id='train_rf', ...)', 'preprocess >> train_group >> select_best' with undefined tasks, 'my_function'), and Airflow 3.x would reject 'schedule_interval=' in favor of 'schedule='. No pseudocode-only sections, so it stays above 3.

4 / 5

Workflow Clarity

The Quick Start gives a clearly numbered 1-5 sequence ending in a triggered pipeline and UI verification, the 7-stage pipeline model is ordered data-to-monitoring, and the Known Issues entries model a validate-and-fail-fast pattern (explicit ValueError raises with fix guidance). Missing an end-of-run validation checkpoint for the quick-start pipeline itself and the stage list lacks explicit per-step checkpoints, so it does not reach 5; the sequence is coherent and mostly checkpointed, above 3.

4 / 5

Progressive Disclosure

A dedicated 'When to Load References' section signals each of the three real, one-level-deep reference files with concrete load conditions (e.g., 'Load references/airflow-patterns.md when building complex DAGs with error handling...'), and all cited files exist. Not a 5 because the SKILL.md body itself carries substantial detail that should live in those references (the seven known-issues walkthroughs and the second full DAG example), making the overview heavier than the anchor's 'concise getting-started content' split.

4 / 5

Total

15

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description with an explicit 'Use for' clause, named tools, and natural trigger phrases covering both use cases and failure modes. The main improvement would be broadening trigger synonyms and slightly sharpening the headline action beyond the generic 'Automate ML workflows'.

DimensionReasoningScore

Specificity

The description names the domain ('ML workflows'), the concrete tooling ('Airflow, Kubeflow, MLflow'), and several specific capabilities ('reproducible pipelines, retraining schedules, experiment tracking'). It falls short of the 5 anchor because the headline action 'Automate ML workflows' is generic and capabilities like model registry/deployment are only implicit in the tool names.

4 / 5

Completeness

Explicitly answers both questions: what ('Automate ML workflows with Airflow, Kubeflow, MLflow') and when via an explicit 'Use for' clause with concrete trigger phrases including failure-mode triggers ('task failures, dependency errors'). This matches the 5-anchor exemplar pattern closely; the 4 anchor would require the 'when' to be less explicit or specific, which it is not.

5 / 5

Trigger Term Quality

Natural phrases users would say are present: 'reproducible pipelines', 'retraining schedules', 'MLOps', 'experiment tracking issues', 'task failures', 'dependency errors'. Not a 5 because common synonyms like 'DAG', 'model training', 'pipeline monitoring', or 'Airflow DAG' are absent from the description itself (they only appear in metadata keywords).

4 / 5

Distinctiveness Conflict Risk

The named orchestration tools (Airflow/Kubeflow/MLflow) carve a clear ML-pipeline niche that is unlikely to trigger for unrelated skills. Not a 5 because generic failure triggers ('task failures, dependency errors') could overlap with general CI/CD or debugging skills outside ML pipelines.

4 / 5

Total

17

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

Total

15

/

16

Passed

Repository
secondsky/claude-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.