CtrlK
BlogDocsLog inGet started
Tessl Logo

deepspeed

Expert guidance for distributed training with DeepSpeed - ZeRO optimization stages, pipeline parallelism, FP16/BF16/FP8, 1-bit Adam, sparse attention

44

Quality

45%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./backend/cli/skills/ml-training/deepspeed/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

25%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

This is the anti-pattern the rubric warns about: a monolithic dump of scraped official documentation pasted directly into SKILL.md, duplicating an existing references/ bundle. It scores slightly above the floor only because the inlined config JSON and API snippets are genuinely executable. It needs a radical cut: a lean quick-start with pointers to the real reference files.

Suggestions

Delete the inlined "Common Patterns" documentation walls and replace them with a concise quick-start (install deepspeed, a minimal working ds_config.json, the `deepspeed --hostfile ... train.py` launch command) — this alone would move conciseness from 1 toward 4-5 since the same material already lives in references/tutorials.md and references/other.md.

Add an explicit multi-step workflow with a validation checkpoint, e.g.: 1) `ds_report` to verify operators/build, 2) write ds_config.json, 3) launch with the deepspeed launcher, 4) check logs/wall_clock_breakdown output — giving the skill the sequence and feedback loop it currently lacks.

Fix reference navigation: list all nine actual files (including index.md), remove pointers to nonexistent files ("getting_started", "api", "guides"), describe what each reference actually contains instead of "08.md - 08 documentation", and delete the empty scripts/ and assets/ placeholder sections and stray one-word code blocks.

DimensionReasoningScore

Conciseness

The body inlines entire scraped documentation pages (DeepNVMe tutorial, MoE tutorial, LRRT, features overview, flops profiler, the full JSON config reference, Monitor) as single-paragraph walls of text — thousands of tokens of material Claude either already knows or that duplicates the references/ bundle. Scraping artifacts ("Updated: November 5, 2025 Previous Next", stray one-word code blocks like `libaio`, `32`, `80`) add pure padding.

1 / 5

Actionability

There is genuinely concrete, executable material — the aio_handle/gds_handle creation snippets, complete ds_config JSON examples (fp16, zero_optimization, optimizers, sparse_attention), get_model_profile usage, and the ds_nvme_tune command. However, it is buried in undifferentiated prose, key steps are missing (no install/launch sequence), and many pointers are dead text ("Refer to the tutorial", "Please see the core API doc") with no actual links.

3 / 5

Workflow Clarity

No real multi-step workflow exists: there is no install → configure → launch sequence, no validation checkpoints (e.g., verifying with ds_report), and "Working with This Skill" offers only vague pointers, two of which ("getting_started", "api/guides" reference files) do not exist in the bundle. Only the trivial "Updating: re-run the scraper" sequence is present.

2 / 5

Progressive Disclosure

Despite a populated references/ bundle (9 files, one level deep) that is listed in a "Reference Files" section, the body inlines the full documentation anyway — the exact content that belongs in those files. Navigation is further weakened by referencing nonexistent files ("getting_started", "api", "guides"), omitting index.md from the listing, and including placeholder sections for scripts/ and assets/ directories that do not exist.

2 / 5

Total

8

/

20

Passed

Description

66%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A reasonably specific, well-scoped description with good domain keywords, but it lacks any explicit trigger ('Use when...') guidance, which caps its completeness. It reads as a feature inventory rather than an action-plus-trigger statement.

Suggestions

Add an explicit trigger clause, e.g., "Use when training or fine-tuning models with DeepSpeed, configuring ZeRO stages or a deepspeed_config.json, or debugging distributed training setups" — this would raise completeness from 3.

Include natural user phrasings and synonyms as trigger terms ("mixed precision", "model/optimizer offloading", "deepspeed_config.json", "zero stage") to improve trigger-term coverage toward 5.

Lead with concrete action verbs ("Configure", "Enable", "Debug") instead of the generic "Expert guidance for" framing.

DimensionReasoningScore

Specificity

Names the domain ("distributed training with DeepSpeed") and enumerates several concrete capability areas ("ZeRO optimization stages, pipeline parallelism, FP16/BF16/FP8, 1-bit Adam, sparse attention"). These are feature topics rather than action verbs, and "Expert guidance" is generic, so it does not reach the comprehensive multiple-concrete-actions level of a 5.

4 / 5

Completeness

The "what" is clear (expert guidance for DeepSpeed distributed training with named features), but there is no "Use when..." clause or equivalent trigger guidance; the "when" is only weakly implied by the feature list. Per the rubric, a missing explicit 'Use when' caps completeness at 3.

3 / 5

Trigger Term Quality

Good natural keyword coverage — "distributed training", "DeepSpeed", "ZeRO", "pipeline parallelism", "FP16/BF16/FP8", "1-bit Adam", "sparse attention" — terms a user would plausibly say when needing this skill. A few natural variations are missing (e.g., "mixed precision", "model offloading", "deepspeed config"), keeping it below 5.

4 / 5

Distinctiveness Conflict Risk

"DeepSpeed" plus DeepSpeed-specific features (ZeRO, 1-bit Adam) establish a clear niche with minimal conflict risk against unrelated skills. Minor overlap risk remains with generic distributed-training / PyTorch-scaling skills, since "distributed training" alone is broad.

4 / 5

Total

15

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

14

/

16

Passed

Repository
synthetic-sciences/openscience
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.