CtrlK
BlogDocsLog inGet started
Tessl Logo

moe-training

Train Mixture of Experts (MoE) models using DeepSpeed or HuggingFace. Use when training large-scale models with limited compute (5× cost reduction vs dense models), implementing sparse architectures like Mixtral 8x7B or DeepSeek-V3, or scaling model capacity without proportional compute increase. Covers MoE architectures, routing mechanisms, load balancing, expert parallelism, and inference optimization.

62

Quality

75%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./backend/cli/skills/ml-training/moe-training/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

63%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A code-dense, largely executable reference with good sequencing and tuning feedback loops, but it under-uses its own bundle: the 500+ line body duplicates content that already lives in the three reference files, and internally repeats DeepSpeed scripts and configs. Trimming the inline duplicates and pushing detail into the existing references would raise both conciseness and progressive disclosure.

Suggestions

Offload the Mixtral 8x7B architecture block, the PR-MoE section, and the Inference Optimization section into the existing references (architectures.md, training.md, inference.md) and replace them with one-line pointers, mirroring the 'See Also' pattern already used.

Remove the internal duplication: the Quick Start DeepSpeed command and the 'Training Script' section repeat nearly identical flag sets, and the Core Concepts 'moe' config block duplicates the one in 'Training Configuration' — keep one of each.

Cut concept re-explanations (the ASCII routing flow diagram, the 'Key Components' bullets, and the 'When to Use This Skill' section that restates the frontmatter description) and replace the last one with a short pointer to the description.

DimensionReasoningScore

Conciseness

Most code blocks are dense and useful, but the body is padded: two DeepSpeed training scripts and two near-duplicate 'moe' config JSON blocks appear inline, the 'When to Use' section repeats the description, and the ASCII routing diagram and 'Key Components' bullets re-explain concepts Claude already knows. Fits 'mostly efficient but could be tightened' rather than the noticeably-verbose anchor 2, since the bulk of the code earns its place.

3 / 5

Actionability

Mostly executable guidance: complete MoELayer and MixtralMoEBlock classes, full DeepSpeed commands, and copy-paste config JSONs. Minor gaps keep it below 5 — the Expert Choice routing snippet is a fragment and moe_inference calls a non-existent 'model.load_expert' API.

4 / 5

Workflow Clarity

The body is logically sequenced (Installation → Quick Start → Core Concepts → Training Configuration → Advanced → Best Practices) with tuning feedback loops ('If load imbalance persists... increase aux loss', 'If training unstable, increase z-loss'). Below 5 because there is no explicit validation or checkpoint guidance for a long training run (e.g. how to verify routing health before committing to 500k iterations).

4 / 5

Progressive Disclosure

References are real, one level deep, and clearly signaled in 'See Also', but substantial content that duplicates the bundle is inlined instead of offloaded: the Mixtral block mirrors references/architectures.md, the PR-MoE and training scripts mirror references/training.md, and the inference section mirrors references/inference.md. This matches 'content that should be separate is inline' rather than anchor 4, where only minor placement gaps exist.

3 / 5

Total

14

/

20

Passed

Description

87%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: concrete domain and tools, explicit 'Use when...' triggers with specific model names, and clear distinctiveness. Minor room to convert topical coverage terms into action verbs and add one or two common synonyms.

DimensionReasoningScore

Specificity

Names the domain and concrete tooling ('Train Mixture of Experts (MoE) models using DeepSpeed or HuggingFace') plus topic areas ('routing mechanisms, load balancing, expert parallelism, and inference optimization'), matching the 'several specific actions; minor gaps' anchor. Not a 5 because the 'Covers...' clause lists topics rather than concrete actions.

4 / 5

Completeness

Explicitly answers both 'what' (train MoE models using DeepSpeed or HuggingFace, covering routing, load balancing, expert parallelism) and 'when' ('Use when training large-scale models with limited compute..., implementing sparse architectures like Mixtral 8x7B or DeepSeek-V3, or scaling model capacity...') with concrete trigger phrases — the anchor 5 example.

5 / 5

Trigger Term Quality

Includes natural terms users would say ('Mixture of Experts', 'MoE', 'Mixtral 8x7B', 'DeepSeek-V3', 'DeepSpeed', 'sparse architectures'), giving good keyword coverage. Not a 5 because a few common variations (e.g. 'sparse models', 'Switch Transformer', 'expert parallel') are missing.

4 / 5

Distinctiveness Conflict Risk

A clear niche (MoE training) with distinct triggers including specific model names (Mixtral 8x7B, DeepSeek-V3), making wrong-skill triggering unlikely; at most negligible overlap with a generic LLM-training skill.

5 / 5

Total

18

/

20

Passed

Validation

75%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 12 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (536 lines); consider splitting into references/ and linking

Warning

metadata_version

'metadata.version' is missing

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

relative_links

Relative link issues: 1 missing

Warning

Total

12

/

16

Passed

Repository
synthetic-sciences/openscience
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.