CtrlK
BlogDocsLog inGet started
Tessl Logo

simpo-training

Simple Preference Optimization for LLM alignment. Reference-free alternative to DPO with better performance (+6.4 points on AlpacaEval 2.0). No reference model needed, more efficient than DPO. Use for preference alignment when want simpler, faster training than DPO/PPO.

63

Quality

76%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./backend/cli/skills/ml-training/simpo/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

82%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A strong, execution-focused body: complete configs for three representative workflows, a useful troubleshooting section with concrete parameter fixes, and well-signaled progressive disclosure into three real reference files. Weaknesses are minor — some duplicated commands, no post-launch validation step, and a few paths that point into the external alignment-handbook repo rather than the skill bundle.

DimensionReasoningScore

Conciseness

Mostly lean config/command blocks with little explanatory padding, but the Quick start repeats Workflow 1's launch command verbatim and the install section pins version-locked notes ('Install PyTorch 2.2.2') that could be trimmed.

4 / 5

Actionability

Fully executable, copy-paste-ready YAML configs and accelerate launch commands covering the common cases (base model, instruct model, and reasoning-intensive training), plus concrete numeric fixes for each troubleshooting scenario.

5 / 5

Workflow Clarity

The install -> write config -> launch sequence is clear and the Common issues section gives explicit error-recovery guidance, but there is no validation checkpoint after launching (e.g. verifying loss curves or running an eval), keeping it below a 5.

4 / 5

Progressive Disclosure

Advanced topics are correctly split into three real, one-level-deep reference files (loss-functions.md, hyperparameters.md, datasets.md), each clearly signaled with its scope. Minor gaps: inline hardware tables and algorithm comparison could live in references, and body commands reference repo paths (scripts/run_simpo.py) not present in the skill bundle, which slightly muddies navigation.

4 / 5

Total

17

/

20

Passed

Description

70%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A solid description that clearly identifies what SimPO is and when to prefer it over DPO/PPO, with a quantified performance claim and an explicit 'Use for' trigger. Main weaknesses are attribute-style framing instead of concrete actions and a slightly garbled trigger clause ('when want simpler, faster training').

Suggestions

Rewrite the trigger clause into fluent natural phrasing, e.g. 'Use when the user wants preference alignment/RLHF-style training and prefers simpler, faster training than DPO or PPO.'

Add commonly-said synonyms and related terms (RLHF, reward model, preference data, chosen/rejected pairs) to improve trigger matching.

Reframe property claims as concrete actions the skill performs (e.g. 'Trains models on chosen/rejected preference pairs without a reference model') to raise specificity.

DimensionReasoningScore

Specificity

Names the domain and a couple of concrete properties ('Reference-free alternative to DPO', 'No reference model needed, more efficient than DPO'), but reads as attribute claims rather than a comprehensive list of concrete actions the skill performs.

3 / 5

Completeness

Clearly answers 'what' (reference-free preference optimization alternative to DPO) and includes an explicit 'Use for preference alignment when want simpler, faster training than DPO/PPO' trigger clause, but the 'when' phrasing is slightly awkward and could be more concrete.

4 / 5

Trigger Term Quality

Good natural keyword coverage ('preference alignment', 'DPO', 'PPO', 'simpler, faster training'), though common variations users might say like 'RLHF', 'reward model', or 'preference data' are missing.

4 / 5

Distinctiveness Conflict Risk

The SimPO niche is clear and explicitly positioned against DPO/PPO/GRPO, with only minor overlap risk against a generic DPO-training skill.

4 / 5

Total

15

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

referenced_paths_exist

Referenced path issues: 3 missing

Warning

Total

13

/

16

Passed

Repository
synthetic-sciences/openscience
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.