CtrlK
BlogDocsLog inGet started
Tessl Logo

ml-engineer

Build production ML systems with PyTorch 2.x, TensorFlow, and modern ML frameworks. Implements model serving, feature engineering, A/B testing, and monitoring. Use PROACTIVELY for ML model deployment, inference optimization, or production ML infrastructure.

51

Quality

64%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/ml-engineer/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

21%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The content reads as a persona/knowledge dump rather than actionable guidance: it lists technologies Claude already knows without executable examples, commands, or validation checkpoints, and inlines material that belongs in separate reference files. It needs concrete code, a real workflow with validation steps, and progressive disclosure into bundle files.

Suggestions

Replace the catalog-of-tools bullet lists with executable code snippets and commands for the highest-value tasks (e.g., a minimal torch.compile serving example, an A/B test setup, a model-registry promotion command).

Tighten the 'Response Approach' into a concrete workflow with explicit validation checkpoints and feedback loops (validate -> fix -> retry) for destructive/batch operations like model deployment and retraining.

Move the long capability and knowledge-base listings into separate reference files under ./references/ and link to them from a concise SKILL.md overview to improve both conciseness and progressive disclosure.

DimensionReasoningScore

Conciseness

The body catalogs dozens of well-known tools and frameworks (PyTorch, TensorFlow, Docker, Kubernetes, scikit-learn) as bullet lists — material Claude already knows — making it noticeably verbose with several padded, low-value sections.

2 / 5

Actionability

No executable code, commands, or concrete examples appear anywhere; the body only enumerates capability categories and tool names ('Distributed training: PyTorch DDP, Horovod, DeepSpeed') and the Instructions section offers generic advice like 'Clarify goals, constraints, and required inputs.'

1 / 5

Workflow Clarity

The 'Response Approach' gives a rough 8-step sequence but the steps are abstract ('Analyze ML requirements,' 'Design ML system architecture') with no commands, no validation checkpoints, and no feedback loops despite the skill covering batch/destructive operations like deployments and retraining.

2 / 5

Progressive Disclosure

Header-based structure exists, but no bundle files are present and the body inlines a large knowledge-base catalog that should live in separate reference files; navigation references are entirely absent.

3 / 5

Total

8

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific, complete, and well-targeted, clearly stating both the capability and the trigger conditions in third person. Its only mild weakness is that the trigger terms are somewhat jargon-heavy rather than mirroring a user's natural phrasing.

DimensionReasoningScore

Specificity

Lists multiple concrete actions across the domain — 'model serving, feature engineering, A/B testing, and monitoring' plus 'ML model deployment, inference optimization, or production ML infrastructure' — giving comprehensive coverage of what the skill does.

5 / 5

Completeness

Explicitly answers both 'what' (building and operating production ML systems) and 'when' via the 'Use PROACTIVELY for ML model deployment, inference optimization, or production ML infrastructure' trigger clause.

5 / 5

Trigger Term Quality

Good coverage of domain keywords ('ML model deployment,' 'inference optimization,' 'model serving,' 'A/B testing') but the terms lean technical; a few natural user phrasings or synonyms (e.g. 'serve my model,' 'optimize inference latency') are missing.

4 / 5

Distinctiveness Conflict Risk

Occupies a clear niche — production ML infrastructure and serving — with distinct triggers that are unlikely to fire for unrelated skills.

5 / 5

Total

19

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

Total

15

/

16

Passed

Repository
rmyndharis/antigravity-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.