CtrlK
BlogDocsLog inGet started
Tessl Logo

spark-training-gotchas

Preflight and diagnose the ten known failure modes for ML training on NVIDIA DGX Spark. Use when a training run on DGX Spark fails to start, OOMs below the 128GB limit, slows down mid-run, or before any multi-hour training job on GB10.

71

Quality

86%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

86%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured diagnostic skill: lean per-gotcha entries, executable triage commands, and clean one-level-deep references to real bundle files. The only friction is mild table/detail duplication and the absence of an explicit feedback loop in the triage workflow.

Suggestions

Collapse the quick-reference table or the per-gotcha SYMPTOM/FIX lines so the summary and detail don't repeat the same wording (boosts conciseness).

Add an explicit triage feedback loop in 'Fast Triage' (e.g., after preflight.sh, map each FAIL/WARN to its G-section and re-run the relevant check after applying the FIX) to reach a 5 on workflow clarity.

Fix the G6 section's mid-word line wrapping so the prose reads continuously rather than as a narrow ragged column.

DimensionReasoningScore

Conciseness

The body is dense and assumes Claude's competence (no padding about what OOM or training is), but the quick-reference table partly duplicates the per-gotcha SYMPTOM/FIX detail, the minor over-explanation that the 4 anchor describes.

4 / 5

Actionability

Copy-paste commands (torch.version.cuda, drop_caches, nvidia-smi --query-gpu=…), an executable assets/preflight.sh, and concrete per-gotcha CHECK/FIX lines cover the common cases fully.

5 / 5

Workflow Clarity

'Fast Triage' gives a cheapest-checks-first sequence and preflight.sh emits a fixed G-number/PASS/FAIL/WARN checkpoint, but the diagnostic flow lacks an explicit validate-then-retry feedback loop, sitting at 4 rather than 5.

4 / 5

Progressive Disclosure

SKILL.md is a concise overview with a quick-reference table; detailed commands live one level deep in the real references/gotcha-checks.md and automation in the real assets/preflight.sh, with CHECK lines clearly signaling each reference.

5 / 5

Total

18

/

20

Passed

Description

87%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, third-person description that clearly states both capability and trigger conditions for a narrow hardware-specific niche. Minor room for a slightly broader action vocabulary keeps specificity and trigger coverage at 4 rather than 5.

DimensionReasoningScore

Specificity

"Preflight and diagnose the ten known failure modes for ML training" names the domain plus two concrete actions and an enumerated scope, matching the 'several specific actions; minor gaps' anchor rather than the fully comprehensive 5.

4 / 5

Completeness

It gives an explicit 'what' (preflight and diagnose the ten failure modes on DGX Spark) and an explicit 'Use when…' clause with concrete triggers, matching the top anchor.

5 / 5

Trigger Term Quality

Phrases like "fails to start", "OOMs below the 128GB limit", "slows down mid-run", and "multi-hour training job on GB10" are natural user language with good coverage, but a few common synonyms are absent so it sits at 4 rather than 5.

4 / 5

Distinctiveness Conflict Risk

The DGX Spark / GB10 niche with hardware-specific triggers (128GB limit, GB10) is highly distinct with minimal overlap risk against generic training skills.

5 / 5

Total

18

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
wshobson/agents
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.