CtrlK
BlogDocsLog inGet started
Tessl Logo

verl-rl-training

Provides guidance for training LLMs with reinforcement learning using verl (Volcano Engine RL). Use when implementing RLHF, GRPO, PPO, or other RL algorithms for LLM post-training at scale with flexible infrastructure backends.

60

Quality

71%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./backend/cli/skills/ml-training/verl/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

63%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, largely executable skill body with three concrete workflows, checklists, and runnable configs. Its two real defects are duplication: ~100 lines of config-reference and troubleshooting content are inlined even though dedicated reference files exist but are never linked, and there are small quality leaks (an empty section heading, stale version pins, an unwired reward function).

Suggestions

Replace the inline 'Configuration Reference' and 'Common Issues and Solutions' sections with one-line pointers to references/api-reference.md and references/troubleshooting.md, which already contain that content in more depth.

Wire the reward function into the GRPO workflow (show the custom_reward_function config path) so Step 2's code is actually used by Step 4's launch command.

Remove or fill the empty 'Multi-Turn Tool Calling' subsection, and move version-sensitive pins (vllm>=0.8.5,<=0.12.0, 'Avoid vLLM 0.7.x') into the troubleshooting reference where they can be maintained.

DimensionReasoningScore

Conciseness

Mostly dense, verl-specific configuration and commands rather than explanation of known concepts, but there is real redundancy: the inline 'Common Issues and Solutions' section (~60 lines) duplicates references/troubleshooting.md, version-sensitive pins like 'pip install vllm>=0.8.5,<=0.12.0' and 'Avoid vLLM 0.7.x' will rot, and the 'Multi-Turn Tool Calling' heading has no body at all. Not anchor 2 since the padding is confined to a few sections, not anchor 4 since the duplication and dead heading are more than 'minor instances'.

3 / 5

Actionability

The guidance is largely executable: a runnable main_ppo command, a copy-paste dataset builder, a complete reward function, and full YAML configs for three workflows. Below anchor 5 because of small gaps — the reward function is never wired into the training config (no reward_function path/registration is shown), and the empty 'Multi-Turn Tool Calling' section promises content it does not deliver.

4 / 5

Workflow Clarity

Workflows are clearly sequenced with prerequisites checklists and a monitoring/validation step ('Check WandB/TensorBoard', 'Verify reward is increasing', 'Run evaluation on held-out test set'), matching anchor 4's 'clear sequence with most checkpoints present'. Not anchor 5 because there are no feedback loops telling Claude what to do when a checkpoint fails — that recovery guidance exists only in the unreferenced troubleshooting file.

4 / 5

Progressive Disclosure

The body is well sectioned, but the two bundle files (references/api-reference.md, references/troubleshooting.md) are never mentioned or linked anywhere in the body, and the inline 'Configuration Reference' and 'Common Issues' sections duplicate their content. This matches anchor 3 exactly — 'references present but not clearly signaled; content that should be separate is inline' — rather than anchor 2, since the document does have substantial internal structure.

3 / 5

Total

14

/

20

Passed

Description

78%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description with an explicit 'Use when...' trigger clause covering natural algorithm keywords (RLHF, GRPO, PPO) and a clearly defined niche. Its main weakness is the generic 'Provides guidance for' framing, which undersells the specific capabilities the skill delivers.

DimensionReasoningScore

Specificity

The description names the domain ("training LLMs with reinforcement learning using verl") and one concrete action area, but 'Provides guidance for' is generic and it does not enumerate multiple specific capabilities (e.g., reward function definition, backend selection, multi-node launch) that the skill actually covers. It sits above anchor 2 ('names the domain but actions are minimal') because it does list concrete algorithm targets, but below anchor 4, which expects several specific actions.

3 / 5

Completeness

It explicitly answers both questions: what ('Provides guidance for training LLMs with reinforcement learning using verl (Volcano Engine RL)') and when ('Use when implementing RLHF, GRPO, PPO, or other RL algorithms for LLM post-training at scale'). The 'when' clause contains concrete trigger phrases, matching the anchor-5 pattern; re-reading anchor 4 ('when could be more explicit') does not fit better since the trigger is fully explicit.

5 / 5

Trigger Term Quality

Natural trigger terms are present and specific: 'RLHF', 'GRPO', 'PPO', 'reinforcement learning', 'LLM post-training'. It is below anchor 5 because common user phrasings like 'fine-tuning with rewards', 'reward model training', or the library name variants users might type are absent, but above anchor 3 because several real algorithm keywords users would say are covered.

4 / 5

Distinctiveness Conflict Risk

Naming verl and its specific algorithms carves out a clear niche with minimal conflict risk against unrelated skills, but trigger terms like 'PPO', 'RLHF', and 'RL training' overlap with closely related fine-tuning libraries (e.g., TRL, which the body itself names as an alternative), so anchor 4 ('mostly distinct; minor overlap risk with closely related skills') fits better than anchor 5.

4 / 5

Total

16

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

14

/

16

Passed

Repository
synthetic-sciences/openscience
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.