CtrlK
BlogDocsLog inGet started
Tessl Logo

trl-fine-tuning

TRL: SFT, DPO, GRPO, RLOO reward modeling for LLM RLHF.

59

Quality

70%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./optional-skills/mlops/training/trl-fine-tuning/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

75%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A strong, example-driven skill: complete executable code for every major TRL method, clear checklisted workflows, sensible use of one-level-deep reference files, and current guidance (the PPO-removal note). The main deductions are small: a missing templates/ file promised twice, a checklist step without body content, and some repeated PPO/DPO duplication.

Suggestions

Add templates/basic_grpo_training.py to the bundle or remove the two references to it (SKILL.md line 289/469 and references/grpo-training.md line 483).

Write the missing Workflow 2 'Step 4: Evaluate alignment' body content, or drop it from that workflow's checklist.

Consolidate the three PPO-removal mentions into one (the Workflow 1 note) and have the other spots point to it, and trim the Quick-start DPO example that duplicates Workflow 2.

DimensionReasoningScore

Conciseness

The body is dominated by lean, executable code with only one sentence of conceptual framing, but has trimmable repetition: the PPO-removal note appears three times (Workflow 1 note, Step 3 intro, and method selection), the Quick-start DPO example duplicates Workflow 2, and trivial comments like '# Load model' add nothing.

4 / 5

Actionability

Nearly all snippets are copy-paste ready with real imports, real dataset names ('trl-lib/Capybara', 'trl-lib/tldr'), concrete hyperparameters, and CLI alternatives. Minor gaps: the RLOO Python block references undefined 'dataset' and 'tokenizer', and the touted 'production-ready script' templates/basic_grpo_training.py does not exist in the bundle.

4 / 5

Workflow Clarity

All three workflows use copy-able checklists with numbered, well-sequenced steps, and the Common issues section provides error-recovery guidance. However, validation is mostly terminal (Step 4 'Evaluate') rather than checkpointed, and Workflow 2's checklist promises 'Step 4: Evaluate alignment' with no corresponding step content in the body.

4 / 5

Progressive Disclosure

Good structure: five topic-specific reference files in references/ are each linked from the body with descriptive one-line pointers, one level deep. The main gap is the broken reference to templates/basic_grpo_training.py, which is promised twice (body and grpo-training.md) but absent from the bundle.

4 / 5

Total

16

/

20

Passed

Description

66%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A terse, concrete, keyword-dense description that covers TRL's full method surface without fluff, but it is purely a 'what' list — no 'Use when' trigger clause and heavy reliance on acronyms rather than natural phrases. Adding an explicit trigger clause would lift it from good to excellent.

Suggestions

Add a 'Use when...' trigger clause, e.g. 'Use when fine-tuning or aligning LLMs with SFT, DPO, GRPO, RLOO, or a reward model, or when the user mentions RLHF, preference alignment, or post-training.'

Spell out one or two natural-phrase synonyms (fine-tuning, preference alignment, HuggingFace) so it matches how users actually phrase requests, not only acronym-fluent ones.

DimensionReasoningScore

Specificity

"TRL: SFT, DPO, GRPO, RLOO reward modeling" names every major post-training method concretely — several specific capabilities, not generic language — but expresses them as an acronym stack with no action verbs ("Extract text..."-style phrasing), so it sits between the 'several specific actions' and 'comprehensive with actions' anchors, noticeably above the midpoint.

4 / 5

Completeness

The 'what' is clear (five concrete method families under TRL), but there is no 'Use when...' clause or any equivalent trigger guidance — the what is only weakly implied by the domain acronym RLHF, which caps completeness at 3 per the rubric guideline.

3 / 5

Trigger Term Quality

"SFT, DPO, GRPO, RLOO reward modeling for LLM RLHF" gives good keyword coverage of exactly what a user doing this work would say, but omits natural variations like 'fine-tuning', 'preference alignment', 'post-training', and 'HuggingFace' that users commonly use.

4 / 5

Distinctiveness Conflict Risk

TRL plus the method acronyms (SFT/DPO/GRPO/RLOO) form a clear niche with distinct triggers, though generic terms like 'reward modeling' and the un-expanded SFT create minor overlap risk with generic fine-tuning or reward-model skills.

4 / 5

Total

15

/

20

Passed

Validation

75%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 12 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

relative_links

Relative link issues: 2 missing

Warning

Total

12

/

16

Passed

Repository
NousResearch/hermes-agent
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.