CtrlK
BlogDocsLog inGet started
Tessl Logo

slime

RL post-training for LLMs with Megatron and SGLang.

48

Quality

54%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./optional-skills/mlops/slime/SKILL.md

The canonical home for this skill is slime-rl-training in Orchestra-Research/AI-Research-SKILLs

SKILL.md
Quality
Evals
Security

Quality

Content

60%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is highly actionable with concrete, executable commands for standard, async, and multi-turn training, and workflows are well sequenced with checklists. Its two real weaknesses are token efficiency and progressive disclosure: roughly half the body duplicates two existing reference files that are never linked from it, so the bundle's structure goes unused.

Suggestions

Replace the inlined "Common Issues and Solutions" section with a two-line summary and a link to references/troubleshooting.md, keeping only the most common issue inline.

Link references/api-reference.md from the Configuration Reference and Data Buffer sections, moving the duplicated argument tables, architecture diagram, and class definitions into that file by reference.

Remove the near-verbatim duplication between the Quick Start launch command and Workflow 1 Step 3 (keep one, point to it from the other).

DimensionReasoningScore

Conciseness

The body is code-dense and mostly free of conceptual padding, but wastes tokens on real redundancy: the Quick Start launch command nearly repeats Workflow 1's Step 3 verbatim, and the "Common Issues" and "Configuration Reference" sections substantially duplicate content that also lives in the reference files.

3 / 5

Actionability

Concrete, copy-paste-ready guidance throughout: full train.py/train_async.py invocations with flags, JSONL data formats, model-script sourcing, the batch-size constraint equation with a worked example, and multi-task evaluation commands. Minor gaps: illustrative code like custom_generate.py calls undefined helpers (generate_single, extract_tool_call, execute_tool) and the RolloutDataSource snippets are schematic.

4 / 5

Workflow Clarity

Three workflows are clearly sequenced with prerequisites checklists, numbered steps, and post-launch monitoring checkpoints (TensorBoard, reward curves, GPU utilization). Not a 5 because there is no error-recovery/feedback loop if training hangs or reward curves diverge, and the async workflow skips verification steps entirely.

4 / 5

Progressive Disclosure

The bundle provides references/api-reference.md and references/troubleshooting.md, but the body never mentions or links either file; instead it inlines their content — the troubleshooting issues, the configuration reference, the data-buffer classes, and even the identical architecture diagram appear in both SKILL.md and the reference files. This matches anchor 2: "content that clearly belongs in separate files is inlined", with the reference files effectively orphaned rather than merely unclearly signaled.

2 / 5

Total

13

/

20

Passed

Description

48%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description names a specific, distinctive domain and technology stack, but reads as a bare noun phrase: it states no actions, includes no 'Use when...' trigger guidance, and omits natural synonyms like 'reinforcement learning' or 'GRPO'. It is well above a generic placeholder but clearly below the good examples, which pair concrete actions with explicit use conditions.

Suggestions

State concrete actions with verbs, e.g. "Trains and post-trains LLMs with reinforcement learning (GRPO/PPO), pairing Megatron-LM training with SGLang rollout generation."

Add an explicit trigger clause, e.g. "Use when fine-tuning or RL-training LLMs (GLM, Qwen, DeepSeek, Llama) with GRPO, or when the user mentions post-training, RLHF, or rollout generation."

Include natural synonyms and variations users would actually say — "reinforcement learning" spelled out, "RLHF", "fine-tuning" — not just the abbreviations "RL" and "post-training".

DimensionReasoningScore

Specificity

"RL post-training for LLMs with Megatron and SGLang" names the domain and two concrete technologies, but contains no verbs or actions at all — fewer concrete actions than the anchor-2 example "Processes PDF files", yet more substantive than the purely abstract anchor 1.

2 / 5

Completeness

A 'what' is identifiable through the specific domain and technology stack, but a 'when should Claude use it' clause is entirely absent, which caps completeness at 3 per the judging guidelines. Not a 2 because the 'what' names concrete technologies rather than being vague.

3 / 5

Trigger Term Quality

Relevant keywords are present ("RL", "post-training", "LLMs", "Megatron", "SGLang") but common variations a user would naturally say are missing — "reinforcement learning" spelled out, "GRPO", "fine-tuning", "RLHF" — matching "some relevant keywords but missing common variations or synonyms".

3 / 5

Distinctiveness Conflict Risk

The Megatron+SGLang pairing carves a clear niche, but adjacent RL-training frameworks (verl and similar) create minor overlap risk, and the absence of any trigger phrase leaves the distinction implicit — matching "mostly distinct; minor overlap risk" rather than the fully distinct anchor 5.

4 / 5

Total

12

/

20

Passed

Validation

75%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 12 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

referenced_paths_exist

Referenced path issues: 4 missing, 2 deeper-than-1-level

Warning

Total

12

/

16

Passed

Repository
NousResearch/hermes-agent
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.