CtrlK
BlogDocsLog inGet started
Tessl Logo

ml-engineer

Build production ML systems with PyTorch 2.x, TensorFlow, and modern ML frameworks. Implements model serving, feature engineering, A/B testing, and monitoring.

44

Quality

44%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./plugins/AI-Agents-Safe-Coding-Skills-claude/skills/ml-engineer/SKILL.md

The canonical home for this skill is ml-engineer in administrakt0r/AI-Agents-Safe-Coding-Skills

SKILL.md
Quality
Evals
Security

Quality

Content

23%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a persona-and-capability catalog: heavily padded with enumerations of ML tools Claude already knows, with abstract directives and no executable code, commands, or validation checkpoints, and its one external reference points to a non-existent file. It needs to be replaced with concrete, actionable guidance and real reference files.

Suggestions

Replace the abstract Instructions/Response-Approach directives with concrete, executable guidance: real code snippets, commands, or copy-paste patterns for the common cases (model serving setup, A/B test scaffolding, feature-store calls).

Delete the Capabilities/Knowledge Base/Behavioral Traits catalogs of tools Claude already knows, or move any genuinely non-obvious detail into real reference files under references/.

Create the referenced `resources/implementation-playbook.md` (and switch to the references/ convention) so the one progressive-disclosure pointer resolves, and add explicit validation checkpoints (e.g. validate → fix → retry) to the production/deployment workflow.

DimensionReasoningScore

Conciseness

Noticeably verbose: ~100 lines of the Capabilities/Knowledge Base sections enumerate frameworks and tools (PyTorch, TensorFlow, Kubernetes, Spark, MLflow, etc.) that Claude already knows, which is padded catalog content; not 1 because it lists tools rather than condescendingly explaining what they are in prose, and not 3 because the padding is substantial rather than occasional.

2 / 5

Actionability

Entirely abstract directives with no concrete code or commands anywhere ("Clarify goals", "Apply relevant best practices", "Analyze ML requirements", "Design ML system architecture"), matching the 1-anchor's "entirely vague or abstract; only describes rather than instructs"; not 2 because there is not even minimal concrete guidance such as a named tool plus a concrete action.

1 / 5

Workflow Clarity

The 8-step Response Approach is a coherent numbered sequence but has no validation checkpoints or feedback loops, and the production/batch-deployment context triggers the cap-at-3 rule for missing validation; not 2 because the sequence is logically ordered rather than gappy, and not 4 because no checkpoints are present.

3 / 5

Progressive Disclosure

The 80-line capability tool catalog and examples are inlined rather than split into files, and the single external reference (`resources/implementation-playbook.md`) points to a file that does not exist; not 1 because real section headers provide structure, and not 3 because the reference is a broken pointer rather than merely "not clearly signaled" and far too much content is inlined for a 160-line skill.

2 / 5

Total

8

/

20

Passed

Description

66%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description states a clear set of concrete ML production capabilities with natural trigger keywords and a reasonably distinct niche, but it omits any explicit "Use when..." trigger guidance, capping completeness at 3. Tightening the action terms and adding a when-to-use clause would lift it toward the top anchors.

Suggestions

Add an explicit "Use when..." clause naming concrete user triggers, e.g. "Use when building or serving production ML models, running A/B tests on model versions, or setting up feature stores."

Sharpen the listed actions from broad categories to specific operations (e.g. "serve models with TorchServe/BentoML", "run A/B tests with multi-armed bandits") to push specificity from 4 to 5.

Include natural synonyms users say (MLOps, inference, fine-tuning, model deployment) to broaden trigger-term coverage.

DimensionReasoningScore

Specificity

Lists several concrete actions ("model serving", "feature engineering", "A/B testing", and "monitoring") with concrete verbs, matching the 4-anchor's "several specific actions; minor gaps"; not 5 because the actions are broad category names rather than sharply concrete operations like "fill forms".

4 / 5

Completeness

Has a clear "what" (build production ML systems, model serving, feature engineering, A/B testing, monitoring) but no "Use when..." clause or equivalent trigger guidance, which per the judging guidelines caps completeness at 3; not 4 because "when" is entirely absent.

3 / 5

Trigger Term Quality

Good keyword coverage of natural terms (PyTorch, TensorFlow, model serving, feature engineering, A/B testing) that users would say; not 5 because synonyms/common variations like MLOps, inference, training, or fine-tuning are missing, and not 3 because coverage is clearly good rather than merely "some relevant keywords."

4 / 5

Distinctiveness Conflict Risk

Mostly distinct niche (production ML with named frameworks and serving/A/B-testing operations) with only minor overlap risk against closely related data-science or LLM skills; not 5 because the broad "ML systems" scope still overlaps adjacent ML skills, and not 3 because named frameworks plus the production/serving focus narrow it beyond "somewhat specific."

4 / 5

Total

15

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
administrakt0r/AI-Agents-Safe-Coding-Skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.