CtrlK
BlogDocsLog inGet started
Tessl Logo

slime-rl-training

Provides guidance for LLM post-training with RL using slime, a Megatron+SGLang framework. Use when training GLM models, implementing custom data generation workflows, or needing tight Megatron-LM integration for RL scaling.

58

Quality

68%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./backend/cli/skills/coding/slime/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

63%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, highly actionable body with concrete commands, checklists, and a useful configuration reference. Its main flaws are progressive-disclosure failures — existing reference files are never linked and their content is duplicated inline, with dead pointers to nonexistent example paths — plus several code skeletons presented as if executable.

Suggestions

Replace the inlined "Architecture Overview" and "Common Issues and Solutions" sections with clearly signaled links to references/api-reference.md and references/troubleshooting.md, keeping only a one-line summary or the top issue inline.

Verify bundle-relative paths before citing them: examples/search-r1/, scripts/models/, and examples/ do not exist in this skill bundle — either include the files or point to the upstream repo URL explicitly.

Mark sketch code (custom_generate.py helpers, CustomRewardModel, buffer_filter) as templates with the interface contract stated, or make it executable, so it is not mistaken for copy-paste-ready code.

DimensionReasoningScore

Conciseness

The body is mostly dense, flag-level guidance (commands, config blocks, tables) with little basic-concept padding, but it carries avoidable weight: the Architecture Overview and the entire "Common Issues and Solutions" section duplicate content that already lives in references/api-reference.md and references/troubleshooting.md, and the "Key Features" section restates points made elsewhere. This lands on "mostly efficient but could be tightened"; it is not level 4 because the duplicated reference material is a clear trimming opportunity, and not level 2 since there is no conceptual over-explanation of things Claude already knows.

3 / 5

Actionability

Most guidance is copy-paste executable: docker install commands, complete train.py invocations with flags, JSONL data format examples, and flag-by-flag configuration reference. It stops short of level 5 because several Python blocks are illustrative skeletons rather than runnable code — custom_generate.py calls undefined helpers (extract_tool_call, execute_tool), buffer_filter calls select_best, and CustomRewardModel uses undefined load_model/tokenize — without explicitly justifying that flexibility.

4 / 5

Workflow Clarity

Workflows 1-3 are clearly sequenced with prerequisites checklists, numbered steps, and monitoring checklists ("Verify reward curves are increasing", "Monitor GPU utilization"). It fits level 4: clear sequence with most checkpoints present, but minor validation gaps — there are no explicit validate-then-continue or error-recovery checkpoints for long batch training runs (e.g. what to check after Step 1 before launching an expensive Step 3), which keeps it below the feedback-loop-rich level 5.

4 / 5

Progressive Disclosure

The body is well-sectioned with headers, but scored against the actual bundle: references/api-reference.md and references/troubleshooting.md exist yet are never mentioned or linked anywhere in SKILL.md, while their content (architecture, troubleshooting) is inlined instead; meanwhile the body points to examples/search-r1/, scripts/models/, and examples/ paths that do not exist in this bundle. This matches "references present but not clearly signaled; content that should be separate is inline"; it is above level 2 because the body itself has strong structure, and below level 4 because the provided reference files are completely unsignaled and bypassed.

3 / 5

Total

14

/

20

Passed

Description

73%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A solid, third-person description with an explicit and specific "Use when..." trigger clause and a well-defined niche. Its main weakness is a generic "provides guidance" framing that undersells the concrete capabilities, and slightly thin synonym coverage for RL training terminology.

Suggestions

Replace "Provides guidance for" with concrete verbs, e.g. "Run GRPO/PPO post-training of LLMs with slime (Megatron-LM + SGLang), including async training and custom data-generation workflows".

Add natural trigger synonyms users would say, such as "reinforcement learning" spelled out, "GRPO", "RLHF/RL post-training", or "fine-tuning GLM/Qwen models".

DimensionReasoningScore

Specificity

"Provides guidance for LLM post-training with RL using slime, a Megatron+SGLang framework" names the domain concretely (RL post-training, Megatron/SGLang) and touches capabilities like "implementing custom data generation workflows", but the primary action verb is generic ("provides guidance") and it does not enumerate concrete operations such as GRPO training or custom reward models. This matches the anchor for naming the domain with 1-2 concrete actions; it is below level 4, which requires several specific listed actions, and above level 2 because the domain and framework are precisely identified rather than minimal.

3 / 5

Completeness

Both parts are present: the "what" (guidance for LLM post-training with RL using slime, a Megatron+SGLang framework) and an explicit "Use when training GLM models, implementing custom data generation workflows, or needing tight Megatron-LM integration" clause with concrete triggers. It sits at level 4 rather than 5 because the "what" remains a general statement of purpose instead of listing the concrete capabilities the level-5 anchor requires; it is well above level 3 since the "when" is fully explicit rather than weakly implied.

4 / 5

Trigger Term Quality

Terms like "LLM post-training", "RL", "slime", "Megatron", "SGLang", "GLM models", and "RL scaling" are phrases users would plausibly say when needing this skill. Coverage is good but misses natural variations such as "reinforcement learning" spelled out, "GRPO", or "fine-tuning"; that gap keeps it below the comprehensive-synonym level 5 while clearly exceeding the thin keyword coverage of level 3.

4 / 5

Distinctiveness Conflict Risk

The description pins a clear niche: the named slime framework, Megatron+SGLang stack, and GLM model training. These triggers are unlikely to fire for unrelated skills, matching the clear-niche/minimal-conflict anchor; nothing pushes it toward the overlap-risk levels below.

5 / 5

Total

16

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

referenced_paths_exist

Referenced path issues: 4 missing, 2 deeper-than-1-level

Warning

Total

13

/

16

Passed

Repository
synthetic-sciences/openscience
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.