CtrlK
BlogDocsLog inGet started
Tessl Logo

axolotl

Axolotl: YAML LLM fine-tuning (LoRA, DPO, GRPO).

45

Quality

48%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./optional-skills/mlops/training/axolotl/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

36%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The skill body is an auto-generated pattern catalog with some genuinely useful snippets (FSDP config, NCCL benchmarking command), undermined by substantial boilerplate padding, non-executable bare-identifier code blocks, and references to files and directories that don't exist. It provides no end-to-end fine-tuning workflow and needs a cleanup pass to remove phantom references and padding.

Suggestions

Remove or fix phantom references: point "For Beginners"/"For Specific Features" at the actual reference files (api.md, dataset-formats.md, other.md) and delete the scripts/ and assets/ sections for directories that don't exist.

Cut boilerplate sections ("What's inside", "When to Use This Skill", "Notes", "Updating") and replace bare-identifier code blocks ("context_parallel_size", "integrations") with real config/usage examples or drop them entirely.

Add a minimal end-to-end workflow — e.g., prepare dataset (see dataset-formats.md) → write/validate YAML config → launch with the axolotl CLI → check training output — so users get a sequenced path instead of disconnected patterns.

DimensionReasoningScore

Conciseness

Several sections are pure generated boilerplate padding: "What's inside" restates the description, "When to Use This Skill" is a generic bullet list repeating "axolotl", and "Resources", "Notes", and "Updating" describe scraper mechanics and empty directories ("scripts/", "assets/") that don't exist. Multiple "code" blocks contain only a bare identifier ("context_parallel_size", "integrations") or a single assignment, adding tokens without information.

2 / 5

Actionability

Guidance is mixed: the FSDP YAML block, the NCCL test command, and the drop_long_seq call are executable, but roughly half the code blocks are bare identifiers ("context_parallel_size", "integrations") or API signatures (cli.cloud.modal_.ModalCloud(config, app=None)) with no imports, usage context, or example values. This matches anchor 3 — some concrete guidance but incomplete, pseudocode-like rather than executable.

3 / 5

Workflow Clarity

There is no actual fine-tuning workflow — no sequence such as prepare dataset, write config, run training, validate output. The "Working with This Skill" section offers only vague navigation pointers, and several of them ("getting_started or tutorials reference files", "guides") reference files that do not exist in references/. This is a rough, broken pointer list rather than a defined sequence with checkpoints.

2 / 5

Progressive Disclosure

The bundle is one level deep and real (references/api.md, dataset-formats.md, other.md, index.md), and the "Reference Files" section signals them. However, the body also directs users to nonexistent files (getting_started, tutorials, guides) and nonexistent directories (scripts/, assets/), which breaks navigation. Structure exists but organization gaps and phantom references keep it below anchor 4.

3 / 5

Total

10

/

20

Passed

Description

61%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is concise and names a distinctive niche with relevant trigger terms, but it is a terse label rather than a full description: it lacks action verbs and, critically, any explicit "Use when..." trigger guidance. Adding a when-to-use clause and expanding the capability list into concrete actions would raise both completeness and specificity.

Suggestions

Add an explicit trigger clause, e.g., "Use when working with Axolotl configs, fine-tuning LLMs with LoRA/QLoRA/DPO/GRPO, or debugging Axolotl training runs."

Convert the label into third-person action statements, e.g., "Configures and runs LLM fine-tuning (LoRA, QLoRA, DPO, GRPO) via YAML configs for 100+ models."

Include the natural synonyms and variations users actually say — "fine-tune", "train", "QLoRA", "KTO", "ORPO", ".yml" — to improve trigger term coverage.

DimensionReasoningScore

Specificity

The description names the domain ("YAML LLM fine-tuning") and concrete techniques ("LoRA, DPO, GRPO"), but uses no action verbs — it reads as a topic label rather than a statement of what the skill does. It sits between anchor 2 (domain named, minimal actions) and anchor 3 (domain plus 1-2 concrete actions), and the listed sub-capabilities are more concrete than anchor 2's example.

3 / 5

Completeness

The "what" is clear (YAML LLM fine-tuning with LoRA/DPO/GRPO), but there is no "Use when..." clause or equivalent explicit trigger guidance, which per the judging guidelines caps completeness at 3. The description never states when Claude should invoke it.

3 / 5

Trigger Term Quality

Good natural keyword coverage for the niche: "Axolotl", "YAML", "LLM fine-tuning", "LoRA", "DPO", "GRPO" — terms a user working with this tool would actually say. A few natural variations are missing (e.g., "fine-tune", "train", "QLoRA", "KTO", "ORPO", ".yml" file extensions), keeping it below anchor 5.

4 / 5

Distinctiveness Conflict Risk

The leading "Axolotl:" prefix establishes a clear niche with minimal conflict risk, but the generic technique keywords ("LLM fine-tuning", "LoRA", "DPO", "GRPO") could also trigger closely related fine-tuning skills for other frameworks. This is minor overlap risk with closely related skills — anchor 4 — rather than the clean niche of anchor 5.

4 / 5

Total

14

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

13

/

16

Passed

Repository
NousResearch/hermes-agent
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.