Best practices for reinforcement learning policy optimization. Use when working on RL agents, PPO, SAC, or reward design.
72
88%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
View guide
Passed
No findings from the security scan
Algorithm selection:
Training recipe:
Evaluation:
Common pitfalls:
e2e23c9
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.