CtrlK
BlogDocsLog inGet started
Tessl Logo

verl-rl-training

Provides guidance for training LLMs with reinforcement learning using verl (Volcano Engine RL). Use when implementing RLHF, GRPO, PPO, or other RL algorithms for LLM post-training at scale with flexible infrastructure backends.

67

Quality

83%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

78%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The content is highly actionable with executable code and clearly sequenced workflows, and is reasonably concise for a complex RL-training skill. Its main weakness is progressive disclosure: a broken reference link and orphaned, unlinked reference files undermine navigation.

Suggestions

Fix the broken reference: either create references/multi-turn.md or remove/correct the 'See [references/multi-turn.md]' link in the Multi-Turn Tool Calling section.

Link the existing bundle files from the body (e.g. point the Configuration Reference and Common Issues sections to references/api-reference.md and references/troubleshooting.md) so they are discoverable.

Add an explicit validate->fix->retry feedback loop inside the training workflows (e.g. check loss/reward, abort and adjust config on divergence) rather than only a post-hoc validation checklist.

DimensionReasoningScore

Conciseness

The body is mostly lean and assumes Claude's RL knowledge (no explanations of what PPO/GRPO are), dominated by executable configs and commands; minor trim opportunities exist in the intro paragraph and the 'When to Use'/'Consider alternatives' navigation sections.

4 / 5

Actionability

Abundant copy-paste-ready content covering common cases: install commands, a complete executable reward function, concrete YAML configs, and launch commands for GRPO, PPO, and Megatron multi-node training.

5 / 5

Workflow Clarity

Workflows 1-3 are clearly sequenced with prerequisites checklists and a final 'Monitor and Validate' checklist, but the batch/distributed training workflows lack an explicit inline validate->fix->retry feedback loop, leaving minor validation gaps below the 5 anchor.

4 / 5

Progressive Disclosure

Section structure is good, but the body links to references/multi-turn.md which does not exist, while the two real bundle files (api-reference.md, troubleshooting.md) are not linked from the body, and a large inline 'Common Issues' section overlaps the unreferenced troubleshooting file.

3 / 5

Total

16

/

20

Passed

Description

87%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, well-structured description that explicitly covers both what the skill does and when to use it, anchored by a named tool and concrete algorithm triggers. The main verb is slightly generic and a few common RL synonyms are absent, but it is clearly above the midpoint on every dimension.

DimensionReasoningScore

Specificity

Names the domain and lists several concrete actions/algorithm targets ('training LLMs with reinforcement learning', 'implementing RLHF, GRPO, PPO... LLM post-training at scale with flexible infrastructure backends'), though the lead verb 'Provides guidance' is slightly generic, keeping it just below the comprehensive 5 anchor.

4 / 5

Completeness

Clearly answers both 'what' ('Provides guidance for training LLMs with reinforcement learning using verl') and 'when' via an explicit 'Use when implementing RLHF, GRPO, PPO...' clause with concrete triggers, matching the 5 anchor.

5 / 5

Trigger Term Quality

Includes natural terms a practitioner would say ('RLHF', 'GRPO', 'PPO', 'reinforcement learning', 'LLM post-training') with good coverage, but is missing common synonyms/variants (e.g. REINFORCE, DPO, 'RL fine-tuning') that would push it to 5.

4 / 5

Distinctiveness Conflict Risk

The named tool (verl / Volcano Engine RL) and specific RL post-training niche with algorithm-specific triggers make it clearly distinct with minimal conflict risk against other skills.

5 / 5

Total

18

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

relative_links

Relative link issues: 1 missing

Warning

referenced_paths_exist

Referenced path issues: 2 missing

Warning

Total

13

/

16

Passed

Repository
synthetic-sciences/openscience
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.