CtrlK
BlogDocsLog inGet started
Tessl Logo

torchtitan

Pretrain LLMs at scale with PyTorch 4D parallelism.

55

Quality

64%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./optional-skills/mlops/torchtitan/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

75%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a strong, execution-focused overview: four clearly sequenced workflows with checklists, concrete multi-node and Float8 guidance, a useful troubleshooting section, and a well-wired reference bundle one level deep. It sits just below top marks due to Quick-start duplication, a non-executable config fragment, and absence of explicit verification checkpoints inside the workflows.

DimensionReasoningScore

Conciseness

The body is dense with commands and configs and avoids explaining concepts Claude already knows, but the Quick start duplicates Workflow 1's commands and the "65%+ speedups" marketing claim adds no actionable value, leaving minor trimmable content that keeps it below a 5.

4 / 5

Actionability

Nearly everything is copy-paste ready (run_train.sh invocations, torchrun, a full SLURM script, an env-var fix, a DCP conversion command), but the model_registry(...) Float8 snippet is an adaptable fragment rather than standalone executable code, matching the "minor gaps" anchor.

4 / 5

Workflow Clarity

All four workflows use checklists with numbered, per-step commands, and a Common issues section provides error recovery for OOM, TP memory, Float8 performance, and checkpoint resharding; however, monitoring/validation appears only in Workflow 1 and there are no explicit verify-early checkpoints (e.g. confirm first checkpoint or loss curve) within the workflows.

4 / 5

Progressive Disclosure

The bundle scores well against the actual structure: all four referenced files (fsdp.md, float8.md, checkpoint.md, custom-models.md) exist, are clearly signaled under Advanced topics, and are one level deep. Minor gaps remain — the full 8B TOML block and the Float8 Python registry detail are inlined in the body where the reference files would be the natural home.

4 / 5

Total

16

/

20

Passed

Description

53%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is concise and names a clear, specific niche (PyTorch-native LLM pretraining with 4D parallelism), but it omits any "when to use" trigger guidance and lacks natural synonyms users would say (distributed training, FSDP, Llama pretraining). It is a serviceable one-liner that would benefit most from an explicit use-when clause and broader trigger-term coverage.

Suggestions

Add an explicit trigger clause, e.g. "Use when pretraining or scaling up LLM training runs, or when the user mentions distributed training, FSDP/FSDP2, tensor/pipeline parallelism, or torchtitan."

Include natural synonyms users would actually say — "distributed training", "FSDP", "Llama pretraining", "torchtitan" — rather than only the technical term "4D parallelism".

Mention one or two more concrete capabilities (e.g. Float8 quantized training on H100s, multi-node SLURM launching) to raise specificity and distinctiveness.

DimensionReasoningScore

Specificity

"Pretrain LLMs at scale with PyTorch 4D parallelism" names the domain and one concrete action with its mechanism, but stops short of listing several specific capabilities (FSDP2/TP/PP/CP, Float8, checkpointing) that would warrant a 4.

3 / 5

Completeness

The "what" is clear (pretrain LLMs with PyTorch 4D parallelism) but there is no "Use when..." clause or equivalent trigger guidance, which caps completeness at 3 per the judging guidelines.

3 / 5

Trigger Term Quality

Relevant keywords like "Pretrain", "LLMs", "PyTorch", and "4D parallelism" are present, but common user phrasings such as "distributed training", "FSDP", "torchtitan", and model names (Llama) are missing, matching the "some relevant keywords but missing common variations" anchor.

3 / 5

Distinctiveness Conflict Risk

"Pretrain LLMs" with "4D parallelism" carves a mostly distinct niche versus fine-tuning or inference skills, though the description never names torchtitan itself and could overlap with adjacent distributed-training skills, keeping it below 5.

4 / 5

Total

13

/

20

Passed

Validation

75%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 12 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

referenced_paths_exist

Referenced path issues: 2 missing

Warning

Total

12

/

16

Passed

Repository
NousResearch/hermes-agent
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.