CtrlK
BlogDocsLog inGet started
Tessl Logo

preference-optimization

Align a fine-tuned model with preference data using DPO, ORPO, KTO, or SimPO. Use when preference pairs or thumbs-up/down feedback exist, when choosing between preference-optimization methods, or when a DPO run needs hyperparameters or debugging.

65

Quality

79%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./plugins/llm-finetuning/skills/preference-optimization/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

71%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured routing/decision skill with genuinely concrete guidance and an exemplary reference split — the bundle file is real, one level deep, and carries the executable configs the body defers. The main drag is repetition (SimPO caveats and the catastrophic-forgetting note each stated multiple times) and a pseudocode pair-construction snippet, which cost it on conciseness and actionability.

Suggestions

State the catastrophic-forgetting note once in the References section — it is currently described in both paragraphs ('plus ... a catastrophic-forgetting note' and 'also carries the catastrophic-forgetting note').

Trim the SimPO length-bias/sweep-budget guidance to the table row plus one prose mention; the bullet and the worked example both restate it, and the worked example can carry the point alone.

Make the pair-construction snippet executable (define the closest() selection or use an explicit index lookup) and cut the code comments that restate the μ−2σ rationale already given in the preceding prose.

DimensionReasoningScore

Conciseness

Mostly efficient and free of beginner-concept padding, but with noticeable redundancy: the catastrophic-forgetting note is stated twice in the References section, SimPO's length-bias/sweep-budget point appears in the table, its bullet, and a worked example, and the μ−2σ rationale is restated almost verbatim in the code comments. This is more than the 'minor instances' of anchor 4, so it sits at anchor 3.

3 / 5

Actionability

Concrete hyperparameters (β=0.1, LR 5e-7–1e-6, 1–2 epochs), a decision table keyed on data shape, worked routing examples, and complete TRL config blocks verified in references/method-configs.md. The pair-construction snippet is pseudocode (closest(...) is undefined), which along with the absence of a copy-paste example for the common KTO/ORPO cases in-body keeps it below anchor 5.

4 / 5

Workflow Clarity

The iterative on-policy DPO pattern is a clearly numbered 4-step loop with an explicit validation checkpoint ('Validate at deployment scale before trusting a ranking') and worked examples disambiguating each routing branch. No destructive/batch-operation cap applies, but the loop lacks error-recovery guidance for a failed round (e.g., what to do when a round degrades the checkpoint), leaving minor validation gaps relative to anchor 5.

4 / 5

Progressive Disclosure

The body is a decision-focused overview — method selection, evidence, production pattern — while the heavy material (full DPOConfig/ORPOConfig/KTOConfig blocks, SimPO sweep grid, Unsloth wrappers) is split into references/method-configs.md, which exists, is one level deep, and is clearly signaled with a description of its contents. Navigation is easy with well-labeled sections and cross-skill pointers.

5 / 5

Total

16

/

20

Passed

Description

87%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: third-person, concise, explicit what-and-when structure with named methods and natural trigger phrases. Its only weakness is modest synonym coverage (no 'RLHF' or 'alignment' terminology) that leaves a small chance of missing users who use those words.

DimensionReasoningScore

Specificity

Names the domain ('align a fine-tuned model with preference data') plus four concrete methods (DPO, ORPO, KTO, SimPO) and support postures (hyperparameters, debugging). It stops short of the 5 anchor's comprehensive multi-action coverage — 'align' is the only verb and downstream outputs like pair construction or config generation are not named — but is well above the 1-2-action coverage of anchor 3.

4 / 5

Completeness

Clearly answers what ('Align a fine-tuned model with preference data using DPO, ORPO, KTO, or SimPO') and when, with an explicit 'Use when' clause listing three concrete trigger conditions. Matches the anchor-5 pattern of both what and when with concrete trigger phrases.

5 / 5

Trigger Term Quality

'preference pairs', 'thumbs-up/down feedback', and 'a DPO run needs hyperparameters or debugging' are natural phrases a user with this need would say. A few common synonyms are missing (e.g., 'RLHF', 'alignment', 'preference tuning'), which keeps it below the comprehensive-synonym coverage of anchor 5.

4 / 5

Distinctiveness Conflict Risk

Occupies a clear niche (preference optimization among fine-tuning methods) with distinct triggers — preference pairs, unpaired thumbs-up/down, DPO debugging — that would not fire for sibling SFT or RL-skill needs. Minimal conflict risk, matching anchor 5.

5 / 5

Total

18

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

referenced_paths_exist

Referenced path issues: 1 missing

Warning

Total

15

/

16

Passed

Repository
wshobson/agents
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.