CtrlK
BlogDocsLog inGet started
Tessl Logo

distributed-llm-pretraining-torchtitan

Provides PyTorch-native distributed LLM pretraining using torchtitan with 4D parallelism (FSDP2, TP, PP, CP). Use when pretraining Llama 3.1, DeepSeek V3, or custom models at scale from 8 to 512+ GPUs with Float8, torch.compile, and distributed checkpointing.

69

Quality

86%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

86%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A high-quality body: executable commands and configs for every workflow, explicit when-to-use vs. alternatives guidance, working one-level-deep references, and an issue-driven troubleshooting section. The remaining gaps are modest — implicit rather than explicit validation steps in the workflows, and a small amount of promotional/duplicated material (speedup claims, checklist headers) that could be trimmed for token efficiency.

Suggestions

Add explicit validation steps to the workflow checklists — e.g., 'Verify loss is decreasing in TensorBoard before step N' or 'Confirm checkpoint files exist in ./outputs/checkpoint after step 4' — to move from implicit to explicit feedback loops.

Trim the 'achieving 65%+ speedups over baselines on H100 GPUs' promotional opener and the per-workflow checklists that restate the step headings that immediately follow, saving tokens without losing information.

Frame the 'Performance benchmarks (H100)' table and the nightly-vs-stable install instructions as time-sensitive (e.g., an 'as of torchtitan 0.2' note or a dedicated versioned section) so version drift does not mislead.

DimensionReasoningScore

Conciseness

The body is dense and largely token-efficient — commands, TOML snippets, and comparison tables with almost no explanation of concepts Claude already knows. Minor trimming opportunities keep it below 5: the promotional opener ('PyTorch's official platform... achieving 65%+ speedups'), checklists that duplicate the immediately-following step headings, and time-sensitive benchmark/version numbers presented outside any 'current as of' framing.

4 / 5

Actionability

Fully executable, copy-paste-ready guidance throughout: install commands, a tokenizer download command, complete TOML config blocks, torchrun/SLURM launch commands, and concrete fixes for each 'Common issues' entry. The examples cover the common cases (single node, multi-node, Float8, 4D parallelism) exactly as the top anchor describes.

5 / 5

Workflow Clarity

Four workflows are clearly sequenced with copy-able checklists, numbered steps, and concrete commands, and the 'Common issues' section provides error-recovery feedback (OOM → activation checkpointing; Float8 slow → filter layers). Validation checkpoints are mostly implicit rather than explicit — e.g., 'Training auto-resumes if checkpoint exists' and 'TensorBoard logs are saved' with no step to verify loss is decreasing or the checkpoint is valid — so it sits below the explicit validate/fix/retry pattern of the 5 anchor and above the missing-checkpoint 3 anchor. These are not destructive or batch operations, so no cap applies.

4 / 5

Progressive Disclosure

SKILL.md is a well-organized overview (quick start, workflows, alternatives, troubleshooting) with a dedicated 'Advanced topics' section that cleanly offloads FSDP2, Float8 recipes, checkpointing, and custom models to four one-level-deep references — all of which exist in ./references/ (fsdp.md, float8.md, checkpoint.md, custom-models.md) — each clearly signaled with descriptive link text, matching the top anchor.

5 / 5

Total

18

/

20

Passed

Description

87%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: third-person, explicit 'what' and 'when' clauses, concrete model names, GPU scale, and technique keywords. Its only weakness is that it enumerates features of a single pretraining action rather than multiple distinct actions, leaving minor gaps versus the strongest specificity and keyword-coverage anchors.

DimensionReasoningScore

Specificity

Names the domain ('PyTorch-native distributed LLM pretraining using torchtitan') and several specific capabilities — '4D parallelism (FSDP2, TP, PP, CP)', 'Float8, torch.compile, and distributed checkpointing' — but it is essentially one primary action (pretraining) with feature modifiers rather than a list of multiple distinct actions, so it falls just below the comprehensive-coverage anchor of 5 and clearly above the 1-2-action anchor of 3.

4 / 5

Completeness

Explicitly answers both: what ('Provides PyTorch-native distributed LLM pretraining using torchtitan with 4D parallelism...') and when ('Use when pretraining Llama 3.1, DeepSeek V3, or custom models at scale from 8 to 512+ GPUs...') with concrete trigger phrases, matching the top anchor.

5 / 5

Trigger Term Quality

Strong natural keywords — 'pretraining', 'Llama 3.1', 'DeepSeek V3', 'at scale from 8 to 512+ GPUs', 'Float8' — that a user with this need would actually say. A few common variations are missing ('training from scratch', 'distributed training', 'fine-tuning vs pretraining' distinction), keeping it below the comprehensive synonym coverage of 5 but comfortably above the 3 anchor.

4 / 5

Distinctiveness Conflict Risk

Clear niche — torchtitan-specific pretraining at 8-512+ GPU scale — with model names and parallelism vocabulary that are unlikely to collide with fine-tuning or inference skills. Written in third person ('Provides...'), which fits the distinct top anchor.

5 / 5

Total

18

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

referenced_paths_exist

Referenced path issues: 2 missing

Warning

Total

13

/

16

Passed

Repository
synthetic-sciences/openscience
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.