Content
78%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
A well-structured, mostly lean skill body with a clear contract, a numbered step-by-step procedure, debug/common-issue checklists, and a complete, verified one-level-deep reference bundle. Its main weaknesses are repetition of the core contract rules across five sections, outline-style rather than fully executable code patterns, and missing embedded validation checkpoints in the main workflow.
Suggestions
Consolidate the repeated contract rules: state each rule once in the Contract section and have the Workflow/Debug/Common-issues checklists reference them briefly instead of restating them, cutting several hundred tokens of duplication.
Replace outline patterns in Steps 2–3 (meta-device build, bottom-up sharding loop) with one short, complete copy-paste code block, or move a full runnable example inline from the references.
Add explicit validation checkpoints in the step sequence, e.g. after init verify WORLD_SIZE and one rank per GPU, and after fully_shard assert parameters are DTensors before constructing the optimizer.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | The body is mostly lean bullet-style instruction with no explanations of concepts Claude already knows, but the five contract rules are restated repeatedly (the optimizer-after-sharding rule appears in the Contract, Step 6, Workflow A, the Debug checklist, and Common issues), which is trimmable duplication — 'efficient; minor instances of over-explanation' rather than the 3 anchor's 'some unnecessary explanation'. | 4 / 5 |
Actionability | Concrete, mostly executable guidance throughout: "torchrun --nproc_per_node", "dist.init_process_group(backend=\"nccl\")", "torch.cuda.set_device(int(os.environ[\"LOCAL_RANK\"]))", "model.to_empty(device=\"cuda\")", and exact policy class arguments. It falls short of a 5 because several patterns are outlines rather than copy-paste code ("if isinstance(m, TransformerBlock): fully_shard(m, ...)" and "with torch.device(\"meta\"): model = ...") and no complete runnable snippet is inline. | 4 / 5 |
Workflow Clarity | Steps are clearly sequenced (0→7) and the Debug checklist plus issue→fix pairs provide error-recovery guidance, but the main procedure lacks embedded validation checkpoints (e.g., verify WORLD_SIZE/all ranks on distinct GPUs after init, verify params are DTensors after sharding) — 'clear sequence with most checkpoints present; minor validation gaps' rather than the explicit validate-then-proceed loop of a 5. | 4 / 5 |
Progressive Disclosure | Every inline "Reference:" path (all 12) exists in references/, is one level deep (verified reference contents contain no nested references), and is topical with a consolidated References section at the end — a clear overview with well-signaled one-level-deep references and easy navigation. | 5 / 5 |
Total | 17 / 20 Passed |