CtrlK
BlogDocsLog inGet started
Tessl Logo

grpo-rlvr-training

Train reasoning and verifiable-task behavior with GRPO and reinforcement learning from verifiable rewards (RLVR). Use when task success is algorithmically checkable (math, code, tool calls, structured output), when designing GRPO reward functions, or when a GRPO run diverges or reward-hacks.

73

Quality

90%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

92%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a well-structured, highly actionable overview: an executable reference recipe, a mandatory reward-inspection gate with a feedback loop, and clean one-level-deep reference files. The only weakness is minor repetition of routing context and reference pointers that could be tightened.

DimensionReasoningScore

Conciseness

The body is efficient — decision rules like "DPO for taste, GRPO for reasoning" and the num_generations floor earn their tokens — but there is minor redundancy: routing context is repeated in the intro and Related skills, and references/reward-functions.md is pointed to three separate times, which could be trimmed.

4 / 5

Actionability

A complete, executable GRPOConfig/GRPOTrainer snippet with settled kwarg values, a concrete success-rate decision tree (never succeeds → SFT; sometimes → proceed), and a failure-mode→variant selection table give copy-paste-ready guidance covering the common cases.

5 / 5

Workflow Clarity

The sequence (routing check → applicability → recipe → inspection gate → variant selection) is explicit, and the 50–100-sample reward inspection is framed as a mandatory gate with a feedback loop ("If the reward function's judgment disagrees... fix the reward function first") before any training run.

5 / 5

Progressive Disclosure

The SKILL.md is a concise overview with two one-level-deep references (references/grpo-memory.md and references/reward-functions.md, both present on disk), each clearly signaled in a References section with a description of its contents, and bulk detail is appropriately split into those files.

5 / 5

Total

19

/

20

Passed

Description

87%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description with a clear what, an explicit multi-clause Use-when trigger set, and a well-delineated GRPO/RLVR niche. Keyword coverage and action specificity are good but leave room for a few more natural synonyms and actions.

DimensionReasoningScore

Specificity

"Train reasoning and verifiable-task behavior with GRPO", "designing GRPO reward functions", and "a GRPO run diverges or reward-hacks" name several concrete actions, but coverage is not fully comprehensive, matching the 'several specific actions; minor gaps' anchor rather than the 5.

4 / 5

Completeness

The description explicitly answers what ("Train reasoning and verifiable-task behavior with GRPO and reinforcement learning from verifiable rewards") and when, with three concrete trigger clauses: "Use when task success is algorithmically checkable..., when designing GRPO reward functions, or when a GRPO run diverges or reward-hacks".

5 / 5

Trigger Term Quality

Natural terms like "GRPO", "RLVR", "reward functions", "diverges", "reward-hacks", and the task domains "math, code, tool calls, structured output" give good keyword coverage; a few natural synonyms (e.g., "reinforcement learning" alone, "verifier") are absent, keeping it below the comprehensive-coverage anchor.

4 / 5

Distinctiveness Conflict Risk

"GRPO", "RLVR", and "verifiable rewards" carve out a clear niche distinct from SFT or preference-optimization skills, and the triggers (reward-function design, divergence, reward hacking) are unlikely to fire for the wrong skill.

5 / 5

Total

18

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
wshobson/agents
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.