CtrlK
BlogDocsLog inGet started
Tessl Logo

mle-workflow

Production machine-learning engineering workflow for data contracts, reproducible training, model evaluation, deployment, monitoring, and rollback. Use when building, reviewing, or hardening ML systems beyond one-off notebooks.

70

Quality

86%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

The canonical home for this skill is mle-workflow in affaan-m/ECC

SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A dense, actionable ML-engineering workflow with executable code, templates, gated validation, and feedback loops throughout. Structure is clear; the only leverage point is that the two big cross-reference tables inflate length and could be split into reference files to improve progressive disclosure and conciseness.

Suggestions

Move the 'Reuse the SWE Surface' and 'Ten MLE Task Simulations' tables into separate reference files (e.g. references/swe-surface.md, references/task-simulations.md) and link to them from SKILL.md, improving both conciseness and progressive_disclosure.

Trim or condense the SWE-surface table rows to the handful of mappings most ML work actually depends on, cutting redundant cross-references that pad the body.

DimensionReasoningScore

Conciseness

Largely lean with no over-explanation of basic ML concepts; templates, code, and checklists earn their tokens, though the two large cross-reference tables (SWE surface, Ten Task Simulations) run long and could be trimmed, fitting the score-4 anchor's 'minor instances that could be trimmed'.

4 / 5

Actionability

Provides fully executable, copy-paste-ready artifacts — the frozen-dataclass config with artifact hashing, the assertion-based promotion gates, and the fill-in templates (Iteration Compact, Observation Ledger) — covering common cases concretely.

5 / 5

Workflow Clarity

The Core Workflow is a clearly numbered 6-step sequence with explicit validation checkpoints and feedback loops (promotion gates that fail closed, Error Analysis Loop 'mistake -> cluster -> hypothesis -> experiment'), plus a closing Review Checklist for complex processes.

5 / 5

Progressive Disclosure

Well-organized single-file skill with clear section headers and easy navigation; no bundle files exist, so everything is one level deep, but the ~340-line monolith with two large reference tables that could live in separate files keeps it just short of the score-5 'clear overview with well-signaled references' anchor.

4 / 5

Total

18

/

20

Passed

Description

85%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description that names concrete capabilities across the ML lifecycle, gives an explicit 'Use when' trigger, and carves out a distinct niche from generic software-engineering skills. The only minor gap is trigger-term breadth (a few synonyms/variations missing).

DimensionReasoningScore

Specificity

Lists multiple concrete actions — 'data contracts, reproducible training, model evaluation, deployment, monitoring, and rollback' — covering the full ML lifecycle comprehensively, matching the score-5 anchor exactly.

5 / 5

Completeness

Explicitly answers both what ('Production machine-learning engineering workflow for...') and when ('Use when building, reviewing, or hardening ML systems beyond one-off notebooks') with concrete trigger phrases, matching the score-5 anchor.

5 / 5

Trigger Term Quality

Includes good natural keywords users would say ('building, reviewing, or hardening ML systems', 'one-off notebooks', 'training', 'model evaluation') but is missing some common variations/synonyms; not quite the comprehensive coverage with synonyms that the score-5 anchor requires.

4 / 5

Distinctiveness Conflict Risk

Targets a clear niche — production ML engineering vs one-off notebooks — with distinct triggers ('hardening ML systems', 'rollback', 'model refresh') and minimal overlap risk with adjacent SWE skills.

5 / 5

Total

19

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
affaan-m/ECC
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.