CtrlK
BlogDocsLog inGet started
Tessl Logo

nlp-alignment

Best practices for LLM alignment techniques including RLHF, DPO, and instruction tuning. Use when working on alignment or safety.

60

Quality

70%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./researchclaw/skills/builtin/domain/nlp-alignment/SKILL.md
SKILL.md
Quality
Evals
Security

LLM Alignment Best Practice

Methods:

  • RLHF: Train reward model → PPO fine-tuning (complex but powerful)
  • DPO: Direct preference optimization (simpler, no reward model needed)
  • GRPO: Group relative policy optimization
  • SFT: Supervised fine-tuning as alignment baseline

Training recipe:

  • Start with SFT on high-quality instruction data
  • DPO: lr=5e-7, beta=0.1, batch_size=64
  • PPO: lr=1e-6, clip=0.2, KL coeff=0.02
  • Use reference model for KL penalty
  • Evaluate on safety benchmarks (TruthfulQA, BBQ, etc.)

Common pitfalls:

  • Reward hacking: model finds shortcuts to high reward
  • Mode collapse: model generates repetitive outputs
  • Catastrophic forgetting: loses general capabilities
Repository
aiming-lab/AutoResearchClaw
Last updated
First committed

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.