CtrlK
BlogDocsLog inGet started
Tessl Logo

ml-engineer

Build production ML systems with PyTorch 2.x, TensorFlow, and modern ML frameworks. Implements model serving, feature engineering, A/B testing, and monitoring.

35

Quality

31%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./plugins/AI-Agents-Safe-Coding-Skills/skills/ml-engineer/SKILL.md

The canonical home for this skill is ml-engineer in administrakt0r/AI-Agents-Safe-Coding-Skills

SKILL.md
Quality
Evals
Security

Quality

Content

3%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

This skill is essentially a persona description or job posting for an ML engineer rather than actionable instructions. It lists hundreds of technologies and concepts Claude already knows without providing any executable code, concrete workflows, or specific guidance. The content would need to be fundamentally restructured from a capabilities catalog into focused, actionable instructions with real code examples and validation steps.

Suggestions

Replace the extensive capabilities lists with 2-3 concrete, executable code examples for the most common tasks (e.g., model serving with TorchServe, setting up a feature store with Feast, implementing A/B testing).

Add explicit multi-step workflows with validation checkpoints for critical operations like model deployment, distributed training setup, and production monitoring configuration.

Remove the 'Behavioral Traits', 'Knowledge Base', 'Example Interactions', and most of the 'Capabilities' bullet lists — these describe what Claude already knows and waste context window tokens.

Create actual bundle files (e.g., resources/implementation-playbook.md, resources/serving-patterns.md) with detailed reference material, and keep SKILL.md as a concise overview pointing to them.

DimensionReasoningScore

Conciseness

Extremely verbose with extensive lists of technologies, frameworks, and capabilities that Claude already knows. The 'Capabilities' section is essentially a resume listing dozens of tools and concepts without adding any actionable knowledge. Behavioral traits, knowledge base, and example interactions are all padding that waste tokens.

1 / 5

Actionability

Contains zero executable code, no concrete commands, no specific examples, and no actual implementation guidance. The entire skill is a description of what an ML engineer does rather than instructions on how to do anything. Even the 'Response Approach' is a vague numbered list of abstract steps.

1 / 5

Workflow Clarity

No concrete workflow is defined. The 'Response Approach' lists 8 abstract steps like 'Analyze ML requirements' and 'Design ML system architecture' with no specifics, no validation checkpoints, and no error recovery. For a skill covering destructive/batch operations like model deployment and distributed training, this is critically insufficient.

1 / 5

Progressive Disclosure

References `resources/implementation-playbook.md` but no bundle file exists to support it. The massive amount of content that should be in separate reference files (capabilities lists, specialized applications, data management) is all inlined in a monolithic format with minimal useful structure.

2 / 5

Total

5

/

20

Passed

Description

58%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description does a reasonable job listing concrete capabilities and naming specific frameworks, giving it decent specificity and distinctiveness. However, it lacks an explicit 'Use when...' clause, which is critical for Claude to know when to select this skill. The trigger terms are adequate but miss common synonyms and natural user phrases like 'machine learning', 'deep learning', or 'MLOps'.

Suggestions

Add an explicit 'Use when...' clause, e.g., 'Use when the user asks about building ML pipelines, deploying models, setting up model serving, or working with PyTorch/TensorFlow in production.'

Include additional natural trigger terms and synonyms such as 'machine learning', 'deep learning', 'MLOps', 'model deployment', 'inference', and 'neural network' to improve discoverability.

Expand the capability list slightly to cover training workflows and data pipelines, which users commonly associate with production ML systems.

DimensionReasoningScore

Specificity

Lists several specific actions (model serving, feature engineering, A/B testing, monitoring) and names concrete frameworks (PyTorch 2.x, TensorFlow). Minor gaps remain—e.g., training pipelines, data preprocessing, or deployment strategies are not mentioned.

4 / 5

Completeness

The 'what' is clearly stated with specific capabilities and frameworks, but there is no explicit 'when' clause or trigger guidance. Per rubric rules, a missing 'Use when...' clause caps completeness at 3.

3 / 5

Trigger Term Quality

Includes relevant keywords like 'PyTorch', 'TensorFlow', 'ML', 'model serving', 'feature engineering', 'A/B testing', and 'monitoring', but misses common user phrases and synonyms such as 'machine learning', 'deep learning', 'inference', 'deployment', 'MLOps', or 'neural network'.

3 / 5

Distinctiveness Conflict Risk

The focus on production ML systems with specific frameworks (PyTorch 2.x, TensorFlow) and production concerns (serving, A/B testing, monitoring) makes it fairly distinct. However, it could overlap with general data science or MLOps skills due to broad terms like 'ML frameworks' and 'monitoring'.

4 / 5

Total

14

/

20

Passed

Validation

90%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation10 / 11 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

10

/

11

Passed

Repository
administrakt0r/AI-Agents-Safe-Coding-Skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.