CtrlK
BlogDocsLog inGet started
Tessl Logo

grpo-rl-training

Expert guidance for GRPO/RL fine-tuning with TRL for reasoning and task-specific model training

52

Quality

60%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./backend/cli/skills/ml-training/grpo-rl-training/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

60%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a strong, largely executable GRPO playbook with a clear workflow, good checklists, and real troubleshooting guidance. Its main weaknesses are a monolithic single-file structure with broken references to non-existent bundle directories, some padding that re-teaches GRPO fundamentals, and a handful of pseudocode gaps (undefined helpers, placeholder trainer calls, one non-existent API call).

Suggestions

Ship the referenced 'templates/' and 'examples/' directories or remove the references to them — currently the body points agents to files that do not exist, which will cause dead-end lookups mid-task.

Split the monolithic file: move the two full GRPOConfig variants, the troubleshooting/pitfalls guide, and the advanced patterns into reference files under references/ and keep SKILL.md as a concise workflow overview, per progressive-disclosure structure.

Cut the 'Core Concepts' GRPO/PPO fundamentals explanation and the 'Mathematical Intuition' pseudocode block (concepts the model already knows), and replace placeholder code with runnable versions by defining the helper functions (extract_final_answer, compute_score) or pointing to a complete example file.

Replace the non-existent 'trainer.generate_completions()' call with a real test-without-training mechanism (e.g., a small script that runs the reward functions against a few sampled completions).

DimensionReasoningScore

Conciseness

Mostly efficient, dense config and code, but includes unnecessary explanation Claude already knows: the 'Core Concepts' section re-explains GRPO/PPO fundamentals with a pseudocode 'Mathematical Intuition' block, and 'Usage Instructions for Agents' plus 'Recommended Reading' restate earlier points or add filler. Could be tightened without losing value.

3 / 5

Actionability

Nearly complete, executable code throughout (reward function templates, two full GRPOConfig variants, GRPOTrainer + LoRA setup, Unsloth path) with only minor gaps: 'GRPOTrainer(model=model, ...)' placeholders, undefined helpers (extract_final_answer, compute_score, CUSTOM_SYSTEM_PROMPT, dataset), and a non-existent 'trainer.generate_completions()' API in the debugging snippet.

4 / 5

Workflow Clarity

A clear 4-step implementation workflow (dataset prep → reward functions → config → training) is supported by before/during/after checklists, healthy-vs-warning training metrics, and a symptom→solution pitfalls table. Not a 5 because validation steps are advisory ('test reward functions on sample data') rather than explicit commands, and the one concrete test path uses a non-existent API.

4 / 5

Progressive Disclosure

The ~560-line body is a single monolithic file inlining content that clearly belongs in separate files (two full config variants, troubleshooting guide, advanced patterns), and it directs agents to 'templates/' and 'examples/' directories that do not exist in the bundle — broken references rather than merely unclear ones.

2 / 5

Total

13

/

20

Passed

Description

61%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description clearly states what the skill does and anchors itself in a distinct, searchable niche (GRPO, TRL, RL fine-tuning), but it entirely lacks a 'when to use' trigger clause and omits several natural trigger synonyms (reinforcement learning, RLHF, reward). Adding an explicit 'Use when...' clause with concrete trigger phrases would resolve its main weakness.

Suggestions

Append a 'Use when...' clause with concrete triggers, e.g., 'Use when the user wants to fine-tune a model with GRPO/TRL, design custom reward functions, or do RL-based training for reasoning or structured output.'

Include common synonyms users actually say — 'reinforcement learning', 'RLHF', 'reward function', 'reward modeling' — so the description matches natural phrasings beyond the GRPO/TRL jargon.

Trim the vague trailing phrase 'for reasoning and task-specific model training' in favor of 2-3 named concrete capabilities (e.g., 'design reward functions, configure GRPOTrainer, monitor reward/KL metrics').

DimensionReasoningScore

Specificity

Names the domain ("GRPO/RL fine-tuning with TRL") and 1-2 actions ("fine-tuning ... for reasoning and task-specific model training"), but lists no broader set of concrete actions, matching the anchor for domain plus 1-2 actions without comprehensive coverage.

3 / 5

Completeness

The 'what' is clearly stated ("Expert guidance for GRPO/RL fine-tuning with TRL"), but there is no 'Use when...' clause or equivalent trigger guidance anywhere, capping completeness at 3 per the judging guidelines.

3 / 5

Trigger Term Quality

Contains the natural terms users would say for this niche ("GRPO", "RL", "fine-tuning", "TRL", "reasoning", "model training"), but misses common variations such as "reinforcement learning" spelled out, "RLHF", and "reward". Good coverage with a few natural terms missing.

4 / 5

Distinctiveness Conflict Risk

"GRPO/RL fine-tuning with TRL" is a mostly distinct niche unlikely to trigger the wrong skill, though the trailing "reasoning and task-specific model training" phrase is broad enough to create minor overlap risk with general SFT/fine-tuning skills.

4 / 5

Total

14

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (582 lines); consider splitting into references/ and linking

Warning

metadata_version

'metadata.version' is missing

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

13

/

16

Passed

Repository
synthetic-sciences/openscience
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.