CtrlK
BlogDocsLog inGet started
Tessl Logo

openrlhf-training

High-performance RLHF framework with Ray+vLLM acceleration. Use for PPO, GRPO, RLOO, DPO training of large models (7B-70B+). Built on Ray, vLLM, ZeRO-3. 2× faster than DeepSpeedChat with distributed architecture and GPU resource sharing.

64

Quality

78%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./backend/cli/skills/ml-training/openrlhf/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

78%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A highly actionable, well-structured skill whose commands are copy-paste ready and whose advanced content is properly split into existing one-level references. Its main flaws are duplicated full command blocks between Quick start and the workflow sections and a redundant marketing-style Performance section, plus missing in-workflow validation checkpoints.

Suggestions

Deduplicate the PPO/GRPO command blocks: show the full command once in Quick start and have Workflow 1/2 reference it with only the differing flags (as the 3-line GRPO delta already does).

Drop or fold the 'Performance' section's marketing claims ('2× faster than DeepSpeedChat') into Hardware requirements, since they add no executable guidance.

Add validation checkpoints to the multi-stage pipeline, e.g. 'check the RM's eval accuracy before launching PPO' and 'confirm reward is increasing in logs within N steps; if not, raise --init_kl_coef per the instability fix below.'

DimensionReasoningScore

Conciseness

The body is command-dense with no concept lecturing, but the full ~25-line PPO command is repeated nearly verbatim in Quick start and Workflow 1, GRPO is shown both as a 3-line delta and again as a full command in Workflow 2, and the 'Performance' section repeats marketing claims — clearly more than minor trimmable material, so below the 'efficient' (4) anchor.

3 / 5

Actionability

Every workflow is a complete, copy-paste-ready command (docker run install, ray job submit for PPO/GRPO, deepspeed for RM and DPO) covering the common cases, with concrete flag-level troubleshooting fixes for OOM, GPU index errors, instability, and slow generation.

5 / 5

Workflow Clarity

Workflow 1 is explicitly sequenced (Step 1 reward model → Step 2 PPO) and the Common issues section covers failure modes, matching 'clear sequence with most checkpoints present'; however the workflows themselves contain no validate-then-proceed checkpoints (e.g. confirming the reward model or training logs before launching the next stage), keeping it below 5.

4 / 5

Progressive Disclosure

The body keeps only commands and decision guidance inline, and all advanced material (hybrid engine, algorithm comparison, multi-node, custom rewards) is pushed to well-signaled, one-level-deep references that all exist as real files, matching the clear-overview anchor exactly.

5 / 5

Total

17

/

20

Passed

Description

78%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description that names its niche precisely with an explicit 'Use for' clause and distinct algorithm-level triggers. Main weaknesses are absent synonyms (RLHF spelled out, post-training, reward model) and marketing claims that displace concrete capability statements.

Suggestions

Add natural synonyms and variations such as 'reinforcement learning from human feedback', 'post-training', and 'RL fine-tuning' to broaden trigger coverage.

Replace marketing claims ('2× faster than DeepSpeedChat', 'High-performance') with one or two more concrete capabilities, e.g. reward model training or colocated GPU sharing.

Make the when-clause more trigger-phrase oriented, e.g. 'Use when the user wants to RL-fine-tune a 7B-70B model with PPO, GRPO, RLOO, or DPO, or needs distributed RLHF training on a Ray cluster.'

DimensionReasoningScore

Specificity

Concrete capabilities are named ("PPO, GRPO, RLOO, DPO training of large models (7B-70B+)"), but "High-performance RLHF framework" and "2× faster than DeepSpeedChat" are marketing-flavored rather than additional concrete actions, so it sits between the several-actions (4) and comprehensive (5) anchors.

4 / 5

Completeness

An explicit when-clause exists ("Use for PPO, GRPO, RLOO, DPO training of large models") and the what is stated, satisfying the 4 anchor; the what is partly buzzword-driven ("High-performance... with distributed architecture and GPU resource sharing") and the when could name more concrete trigger phrases, keeping it below 5.

4 / 5

Trigger Term Quality

Natural terms users would say are present ("RLHF", "PPO, GRPO, RLOO, DPO", "Ray", "vLLM", "large models (7B-70B+)"), but common variations like "reinforcement learning from human feedback", "post-training", or "reward model" are missing, matching the good-but-incomplete keyword anchor.

4 / 5

Distinctiveness Conflict Risk

The combination of named RL algorithms (PPO/GRPO/RLOO/DPO), the Ray/vLLM stack, and the 7B-70B scale creates a clear niche with distinct triggers and minimal risk of firing for unrelated skills.

5 / 5

Total

17

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

14

/

16

Passed

Repository
synthetic-sciences/openscience
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.