CtrlK
BlogDocsLog inGet started
Tessl Logo

mlflow

Track ML experiments, manage model registry with versioning, deploy models to production, and reproduce experiments with MLflow - framework-agnostic ML lifecycle platform

55

Quality

64%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./backend/cli/skills/ml-training/mlflow/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

57%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is highly actionable with correct, comprehensive MLflow examples, but it is verbose: ~700 lines whose Core Concepts, registry, and deployment sections largely duplicate the bundled reference files. There is also no step sequencing or validation guidance for production-impacting operations like stage transitions and deployment.

Suggestions

Cut the body to a lean overview (installation, quick start, pointers) and move the detailed Core Concepts, registry, search, and deployment sections into the existing references/*.md files, linking them contextually at each section instead of only in a final 'See Also' list — this addresses both conciseness and progressive disclosure.

Add a canonical workflow with checkpoints for the production path, e.g. log model → verify the run in the UI/client → register → validate on Staging (test predictions) → only then transition to Production, covering the currently missing feedback loop for risky registry and deployment operations.

Remove promotional filler ('Users: 20,000+ organizations | GitHub Stars: 23k+') and eliminate duplicated content between Quick Start and Core Concepts (logging) and the two autologging sections.

DimensionReasoningScore

Conciseness

The body is code-dense rather than prose-padded, but includes marketing filler ('Users: 20,000+ organizations | GitHub Stars: 23k+'), duplicates parameter/metric logging between Quick Start and Core Concepts, and shows autologging twice — fitting the mostly-efficient-with-unnecessary-excess anchor at 3 rather than the several-padded-sections anchor at 2.

3 / 5

Actionability

Concrete, correct MLflow calls and CLI commands throughout (log_param(s), start_run, register_model, transition_model_version_stage, mlflow models serve with a curl test), but placeholders like train_model(), undefined X_train/y_train, and get_git_commit() leave minor gaps that keep it below the fully copy-paste-ready anchor at 5.

4 / 5

Workflow Clarity

Content is organized by topic rather than a sequenced workflow, and validation checkpoints are largely absent for risky operations — promoting to Production with archive_existing_versions=True and cloud deployments proceed with no verify step, while the curl test after local serving is the only checkpoint. This matches the validation-gaps anchor at 3 and the scoring-note cap for missing feedback loops.

3 / 5

Progressive Disclosure

Three real one-level-deep reference files exist and are described, but only in a terminal 'See Also' section rather than signaled contextually, and the ~700-line body inlines full tracking, registry, and deployment tutorials that duplicate those reference files — matching the structure-present-but-inline-content anchor at 3.

3 / 5

Total

13

/

20

Passed

Description

71%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A specific, tool-anchored description that clearly states what the skill does, but it omits any explicit 'use when' trigger guidance, which both caps completeness and weakens its trigger-term coverage. Adding a trigger clause naming natural user phrases would lift the two weakest dimensions.

Suggestions

Append an explicit trigger clause, e.g. 'Use when the user mentions MLflow, experiment tracking, logging metrics or parameters, or a model registry/production deployment' — this addresses the missing 'when' that caps completeness at 3.

Work natural user phrases like 'log metrics', 'log parameters', 'experiment tracking', and 'model versioning' into the description to improve trigger-term coverage beyond the current four verbs.

Clarify when this skill applies versus adjacent ML tracking tools (e.g., mention it is the MLflow-specific skill) to reduce overlap risk with generic experiment-tracking skills.

DimensionReasoningScore

Specificity

The description lists four concrete actions — 'Track ML experiments', 'manage model registry with versioning', 'deploy models to production', 'reproduce experiments' — comprehensively covering MLflow's core capabilities, matching the multiple-specific-actions anchor rather than the minor-gaps anchor at 4.

5 / 5

Completeness

The 'what' is clearly answered (track, registry, deploy, reproduce), but there is no 'Use when...' clause or equivalent explicit trigger guidance, which per the judging guidelines caps completeness at 3.

3 / 5

Trigger Term Quality

Natural user phrases like 'ML experiments', 'model registry', 'deploy models to production', and the tool name 'MLflow' are present, but common variations users would actually say ('log metrics', 'log parameters', 'experiment tracking', 'artifacts') are missing, fitting the good-coverage-with-a-few-gaps anchor.

4 / 5

Distinctiveness Conflict Risk

Naming MLflow and its lifecycle niche makes it mostly distinct, but without a when-clause, generic phrases like 'track ML experiments' could also match competing tracking skills (e.g., Weights & Biases), so it does not reach the minimal-conflict anchor at 5.

4 / 5

Total

16

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (714 lines); consider splitting into references/ and linking

Warning

metadata_version

'metadata.version' is missing

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

13

/

16

Passed

Repository
synthetic-sciences/openscience
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.