CtrlK
BlogDocsLog inGet started
Tessl Logo

veomni-debug

Use this skill for ANY bug, error, crash, wrong output, loss divergence, gradient explosion, test failure, CUDA error, distributed training hang, checkpoint load failure, or unexpected behavior. Covers both quick fixes (clear root cause) and complex debugging (unclear cause). Trigger: 'fix bug', 'fix error', 'broken', 'crash', 'doesn't work', 'fails with', 'loss NaN', 'training hangs', 'FSDP error', 'OOM'.

71

Quality

89%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

85%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An exceptionally actionable, well-sequenced debugging protocol with strong validation gates and feedback loops throughout. Its weaknesses are mild verbosity in the Phase 2 bisect instructions and a fully inlined structure where the advanced bisect procedure and subagent prompt template would be better served as one-level-deep reference files.

Suggestions

Move the multi-environment dependency-bisect procedure (Phase 2 step 5) into a references/ file (e.g., BISECT.md) and keep a 3-4 line summary with the pointer in SKILL.md, improving both conciseness and progressive disclosure.

Move the verification-subagent prompt template in the Appendix to a reference file and retain a one-line invocation summary with the path.

Tighten the repeated 'uv pip freeze / diff / reconcile non-target differences' instructions, which appear nearly verbatim for both the venv and worktree variants, into a single stated procedure.

DimensionReasoningScore

Conciseness

The body is dominated by terse, load-bearing directives ("Don't skim. Extract 2-3 keywords", "If you can't reproduce, you don't understand it") and executable commands, with only a few passages that could be trimmed (e.g., the repeated freeze-compare/reconcile instructions in Phase 2). Most explanation is non-obvious project-specific knowledge (the '--active is load-bearing' caveat, worktree-vs-venv isolation), which Claude would not already know, so it earns its tokens.

4 / 5

Actionability

Guidance is fully executable throughout: copy-paste-ready uv venv/sync/pip/freeze commands, `git worktree add` invocations, `make patchgen`, `make quality`, `git log --oneline -10`, pytest paths, and exact repo file paths (`.agents/knowledge/constraints.md`, `veomni/distributed/parallel_plan.py`). Placeholders like `<reproducer>` and `<other-version>` are appropriate parameterization rather than pseudocode.

5 / 5

Workflow Clarity

The routing table, phased protocol with todo tracking, 15-minute Quick-Path cutoff, explicit Phase 3 verification gate, verify steps in Phase 4 and the Quick Path, stop conditions with self-sabotage phrases, and the 3-attempt escalation form a clear sequence with explicit validation checkpoints and feedback loops. Not below 5 because every phase has a concrete check and recovery path.

5 / 5

Progressive Disclosure

There are no bundle files and no one-level-deep references from SKILL.md — everything, including the ~60-line dependency-bisect procedure, the verification-subagent prompt template, and domain checklists, is inlined in a 218-line file. Structure and headers are good and the knowledge-file pointers are clear, but content that clearly belongs in a separate reference file (the bisect recipe, the appendix prompt) is inlined, matching the 3 anchor; it is not 4 because the split would genuinely improve navigation, and not 2 because the inline content is well-sectioned and self-navigable.

3 / 5

Total

17

/

20

Passed

Description

87%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description with excellent explicit trigger coverage and a clear what/when structure. Its main weakness is the maximalist 'ANY bug ... or unexpected behavior' scope claim, which creates overlap risk with generic debugging skills despite the strong distributed-training-specific triggers.

Suggestions

Narrow the opening scope claim from 'ANY bug, error, crash ... or unexpected behavior' to the debugging domain this skill actually owns, or lead with the VeOmni/distributed-training context so it does not compete with generic code-fixing skills.

State the concrete actions the skill performs (e.g., 'diagnose, bisect dependency versions, and fix') rather than only naming the covered conditions, to lift specificity from symptom enumeration to concrete capability claims.

DimensionReasoningScore

Specificity

The description enumerates concrete, domain-specific conditions ("loss divergence, gradient explosion, test failure, CUDA error, distributed training hang, checkpoint load failure") and distinguishes scope ("quick fixes (clear root cause) and complex debugging (unclear cause)"), though the actions themselves ("fixes", "debugging") are generic rather than multiple distinct concrete actions.

4 / 5

Completeness

It explicitly answers both questions: what ("Covers both quick fixes (clear root cause) and complex debugging (unclear cause)") and when ("Trigger: 'fix bug', 'fix error', 'broken', ...") with concrete trigger phrases, mirroring the 5-anchor example structure.

5 / 5

Trigger Term Quality

The explicit trigger list covers comprehensive natural phrasing users would actually say — "'fix bug', 'fix error', 'broken', 'crash', 'doesn't work', 'fails with'" — plus domain synonyms like "'loss NaN', 'training hangs', 'FSDP error', 'OOM'", matching the top anchor's coverage including synonyms and domain-specific variants.

5 / 5

Distinctiveness Conflict Risk

Opening with "Use this skill for ANY bug, error, crash, ... or unexpected behavior" is extremely broad and would overlap with any general debugging or code-fixing skill, though domain terms (loss divergence, FSDP, distributed training hang) anchor it to ML training and reduce confusion somewhat. It is somewhat specific but could still overlap with similar generic debugging skills, matching the 3 anchor; it is not 4 because the 'ANY bug' claim invites conflicts, and not 2 because the training-domain triggers give it a real niche.

3 / 5

Total

17

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
ByteDance-Seed/VeOmni
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.