CtrlK
BlogDocsLog inGet started
Tessl Logo

simpo

Reference-free preference alignment, simpler than DPO.

56

Quality

67%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./optional-skills/mlops/simpo/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

86%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The content is a strong, practitioner-oriented skill: fully executable configs and commands, useful decision guidance versus alternatives, and a troubleshooting section that doubles as error recovery. Its only weaknesses are minor redundancy and the absence of explicit post-training validation steps.

DimensionReasoningScore

Conciseness

The body is dense with executable configs and commands and avoids explaining concepts Claude already knows; the only padding is the one-line intro that repeats the description and the duplicated accelerate launch command in Quick start and Workflow 1. Not a 5 because those two redundant elements could be trimmed, but well above the 'noticeably verbose' bar of 2.

4 / 5

Actionability

Fully copy-paste-ready content: complete install commands, full YAML training configs with hyperparameter comments and valid ranges ('beta: 2.0 # Reward scaling (2.0-10.0)'), and exact accelerate launch commands covering the common cases. Not a 4 because Workflow 3's missing launch command is a negligible gap — the command pattern is established twice already.

5 / 5

Workflow Clarity

The install → config → launch sequence is clear and the 'Common issues' section provides symptom → fix recovery guidance (loss divergence, forgetting capabilities, OOM). Not a 5 because there are no explicit validation checkpoints such as verifying training loss or running an evaluation after training; not a 3 because error-recovery guidance and concrete sequence are present.

4 / 5

Progressive Disclosure

The SKILL.md body stays a concise overview with three well-signaled one-level-deep references ([references/loss-functions.md], [references/hyperparameters.md], [references/datasets.md]), all of which exist in the bundle. Not a 4 because the split is exactly right: quick-start and common workflows inline, deep detail (math formulations, tuning guides, dataset formats) delegated to referenced files.

5 / 5

Total

18

/

20

Passed

Description

48%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is concise and technically accurate, naming a clear niche, but it functions more as a tagline than a skill description. It omits any 'when to use' trigger guidance and concrete actions, which significantly limits its discoverability and routing value.

Suggestions

Add an explicit trigger clause, e.g. 'Use when the user wants to train or align an LLM with preference data (chosen/rejected pairs) without a reference model, or mentions SimPO, DPO alternatives, or preference optimization.'

State concrete actions in third person, e.g. 'Trains LLMs with preference data using the SimPO algorithm; includes ready-to-run configs for Mistral, Llama 3, and math/reasoning models.'

Include natural synonyms users would actually say — 'SimPO', 'RLHF', 'preference data', 'fine-tuning' — to improve trigger-term coverage beyond the current three jargon terms.

DimensionReasoningScore

Specificity

The description names the domain ("preference alignment") and a distinguishing property ("Reference-free", "simpler than DPO"), but contains no concrete action verbs like train, fine-tune, or optimize — matching the anchor 'Names the domain but actions are minimal or generic'. It is not a 3 because no concrete actionable capability is listed, only what the method is.

2 / 5

Completeness

It has a reasonably clear 'what' ("Reference-free preference alignment, simpler than DPO") but no 'Use when...' or equivalent trigger guidance, which caps completeness at 3 per the judging guidelines. Not a 4 because 'when' is entirely absent rather than just imprecise.

3 / 5

Trigger Term Quality

Relevant keywords exist ("preference alignment", "reference-free", "DPO") that an ML practitioner would say, but common variations like SimPO, RLHF, fine-tuning, preference optimization, or preference data are missing. It is not a 4 because keyword coverage is thin — three jargon terms with no synonyms or natural user phrasings.

3 / 5

Distinctiveness Conflict Risk

The niche (reference-free preference alignment as a DPO alternative) is fairly distinct, but explicitly invoking "DPO" creates minor overlap risk if a separate DPO skill exists and a user asks about DPO training. It is not a 5 because it lacks distinct trigger phrases that would cleanly route users to it over adjacent alignment skills.

4 / 5

Total

12

/

20

Passed

Validation

75%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 12 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

referenced_paths_exist

Referenced path issues: 3 missing

Warning

Total

12

/

16

Passed

Repository
NousResearch/hermes-agent
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.