CtrlK
BlogDocsLog inGet started
Tessl Logo

pytorch-fsdp2

Adds PyTorch FSDP2 (fully_shard) to training scripts with correct init, sharding, mixed precision/offload config, and distributed checkpointing. Use when models exceed single-GPU memory or when you need DTensor-based sharding with DeviceMesh.

68

Quality

83%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

78%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, mostly lean skill body with a clear contract, a numbered step-by-step procedure, debug/common-issue checklists, and a complete, verified one-level-deep reference bundle. Its main weaknesses are repetition of the core contract rules across five sections, outline-style rather than fully executable code patterns, and missing embedded validation checkpoints in the main workflow.

Suggestions

Consolidate the repeated contract rules: state each rule once in the Contract section and have the Workflow/Debug/Common-issues checklists reference them briefly instead of restating them, cutting several hundred tokens of duplication.

Replace outline patterns in Steps 2–3 (meta-device build, bottom-up sharding loop) with one short, complete copy-paste code block, or move a full runnable example inline from the references.

Add explicit validation checkpoints in the step sequence, e.g. after init verify WORLD_SIZE and one rank per GPU, and after fully_shard assert parameters are DTensors before constructing the optimizer.

DimensionReasoningScore

Conciseness

The body is mostly lean bullet-style instruction with no explanations of concepts Claude already knows, but the five contract rules are restated repeatedly (the optimizer-after-sharding rule appears in the Contract, Step 6, Workflow A, the Debug checklist, and Common issues), which is trimmable duplication — 'efficient; minor instances of over-explanation' rather than the 3 anchor's 'some unnecessary explanation'.

4 / 5

Actionability

Concrete, mostly executable guidance throughout: "torchrun --nproc_per_node", "dist.init_process_group(backend=\"nccl\")", "torch.cuda.set_device(int(os.environ[\"LOCAL_RANK\"]))", "model.to_empty(device=\"cuda\")", and exact policy class arguments. It falls short of a 5 because several patterns are outlines rather than copy-paste code ("if isinstance(m, TransformerBlock): fully_shard(m, ...)" and "with torch.device(\"meta\"): model = ...") and no complete runnable snippet is inline.

4 / 5

Workflow Clarity

Steps are clearly sequenced (0→7) and the Debug checklist plus issue→fix pairs provide error-recovery guidance, but the main procedure lacks embedded validation checkpoints (e.g., verify WORLD_SIZE/all ranks on distinct GPUs after init, verify params are DTensors after sharding) — 'clear sequence with most checkpoints present; minor validation gaps' rather than the explicit validate-then-proceed loop of a 5.

4 / 5

Progressive Disclosure

Every inline "Reference:" path (all 12) exists in references/, is one level deep (verified reference contents contain no nested references), and is topical with a consolidated References section at the end — a clear overview with well-signaled one-level-deep references and easy navigation.

5 / 5

Total

17

/

20

Passed

Description

88%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: it names concrete capabilities comprehensively and has an explicit, concrete 'Use when...' trigger clause with domain-accurate technical terms. The only gaps are a few natural trigger variations (e.g., "FSDP", "multi-GPU") and minor collision risk with closely related FSDP1/DDP skills.

DimensionReasoningScore

Specificity

"Adds PyTorch FSDP2 (fully_shard) to training scripts with correct init, sharding, mixed precision/offload config, and distributed checkpointing" lists multiple specific concrete actions covering the full FSDP2 surface (init, sharding, precision, offload, checkpointing) with no coverage gaps, matching the comprehensive-coverage anchor; it exceeds the 4 anchor, which allows minor gaps.

5 / 5

Completeness

Explicitly answers both questions: the first sentence states what it does (add FSDP2 to training scripts with init, sharding, mixed precision/offload, checkpointing) and "Use when models exceed single-GPU memory or when you need DTensor-based sharding with DeviceMesh" gives concrete when-triggers, mirroring the 5 anchor example; a 4 would require a less explicit 'when'.

5 / 5

Trigger Term Quality

Natural terms users would say are present ("PyTorch FSDP2", "fully_shard", "training scripts", "DTensor", "DeviceMesh"), but common variations users naturally say are missing — e.g. "FSDP" alone, "multi-GPU", "sharded data parallel", "model parallel" — so it fits 'good keyword coverage; a few natural terms missing' rather than the comprehensive synonym/extension coverage of a 5.

4 / 5

Distinctiveness Conflict Risk

The FSDP2/fully_shard/DTensor/DeviceMesh niche is well differentiated from generic training skills, but a user asking for "FSDP" or "sharded data parallel" generically could be routed here when an FSDP1 or DDP skill applies, giving minor overlap risk with closely related skills — the 4 anchor — rather than minimal conflict risk.

4 / 5

Total

18

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
OpenLAIR/dr-claw
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.