CtrlK
BlogDocsLog inGet started
Tessl Logo

research

Conduct post-training research for LLMs using the Tinker API — replicate paper results, explore new training ideas, run and monitor experiments, and document findings. Use this skill whenever the user wants to do research, replicate experiments from a paper or repo, investigate training hypotheses, run experiment sweeps, explore post-training techniques (SFT, RL, DPO, distillation, etc.), set up training, write training code, choose a model, tune hyperparameters, manage checkpoints, export weights, or analyze training logs — even if they just say "try this idea" or "let's see what happens if...".

72

Quality

91%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

85%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An excellent operational skill body: fully executable code for every training approach, a clearly sequenced research methodology with real validation checkpoints, and clean progressive disclosure to ten real reference files. The one weakness is verbosity — a motivational 'researcher mindset' layer and a few concept explanations (GRPO mechanics, curiosity exhortations) pad the document and could be cut without losing actionable content.

Suggestions

Trim the 'You are a researcher... mindset' preamble and the scattered exhortations ('Stay curious between experiments', 'A researcher who doesn't read the literature wastes time') into 2-3 imperative bullets; they restate attitudes Claude already applies and cost ~30 lines.

Drop or compress the conceptual aside '**How GRPO works:** For each problem, the model generates group_size responses...' — Claude knows GRPO; keep only the cookbook-specific facts (group_size parameter, built-in builders).

The benchmark table and model-type table are useful, but the per-section 'Existing recipes:' lines duplicate the 'Code references' section; consolidate into one index to remove repeated listings.

DimensionReasoningScore

Conciseness

The bulk is dense and actionable, but there are several padded passages: the mindset preamble ("You are a researcher. This is not a tool you invoke and forget — it is a mindset that shapes everything you do"), exhortations like "A researcher who doesn't read the literature wastes time rediscovering known results" and "Stay curious between experiments. When results surprise you, dig into why", plus an explanation of how GRPO works conceptually — knowledge Claude already has. This matches anchor 3 ('mostly efficient but includes some unnecessary explanation or could be tightened'); it is not anchor 2 because there is no section of generic concept explanation, and not anchor 4 because the motivational material recurs across multiple sections rather than being a couple of trimmable lines.

3 / 5

Actionability

Nearly every section carries copy-paste-ready, complete code: full chz-blueprint SFT/RL/DPO/distillation configs, benchmark registration code, checkpoint save/resume calls, weight export CLI, and environment setup commands. This is the anchor-5 fit ('fully executable; copy-paste ready code or commands; specific examples cover the common cases'); anchor 4 implies minor gaps in executability that are not present.

5 / 5

Workflow Clarity

A seven-step methodology is explicitly sequenced (understand → know models → set up eval FIRST → prepare data → plan → run/monitor → document), with concrete validation checkpoints woven in: "Verify config before launching", the 4-point data inspection checklist ("decode them back to text and read them"), "Start small, scale up" (tiny → right-model → full), and "Immediately after launch: Confirm the process started... If anything is off, investigate now". This matches anchor 5 (clear sequence, explicit validation, feedback loops for error recovery); the operations involved are training runs rather than destructive batch ops, so no cap applies.

5 / 5

Progressive Disclosure

The body is an overview that consistently defers depth to one-level-deep, well-signaled pointers ("For the full environment protocol... read references/rl.md", "For the complete SDK API reference, read references/sdk.md"), and all 10 files listed in the final "Reference files" section exist in ./references/. This matches anchor 5 ('clear overview with well-signaled one-level-deep references; content appropriately split; easy navigation'); anchor 4's 'minor organization gaps' is not in evidence since every section maps to an existing reference file.

5 / 5

Total

18

/

20

Passed

Description

96%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: third-person, concrete about what the skill does, and explicit about when to use it, with unusually good coverage of informal user phrasings. The only weakness is a very broad when-clause that includes generic terms (research, model choice, training setup) which could fire in contexts where a different training-related skill is more appropriate.

DimensionReasoningScore

Specificity

The description lists multiple concrete actions — "replicate paper results, explore new training ideas, run and monitor experiments, and document findings" plus downstream actions like "tune hyperparameters, manage checkpoints, export weights, or analyze training logs" — covering the full research loop. This matches the anchor for multiple specific concrete actions with comprehensive coverage; nothing above 5 exists and it is well beyond the 4 anchor's 'minor gaps'.

5 / 5

Completeness

Both parts are explicit: the 'what' is the opening sentence of concrete capabilities, and the 'when' is a literal "Use this skill whenever the user wants to..." clause with concrete trigger phrases including quoted informal variants. This is exactly the anchor-5 example structure ('clearly and explicitly answers both what AND when with concrete trigger phrases'); anchor 4 requires the 'when' to be less explicit, which it is not.

5 / 5

Trigger Term Quality

Natural trigger phrases users would actually say are comprehensively covered: "replicate experiments from a paper or repo", "run experiment sweeps", "write training code", "tune hyperparameters", plus technique names (SFT, RL, DPO, distillation) and even informal phrasings ("try this idea", "let's see what happens if..."). This is the comprehensive-coverage-with-synonyms anchor; anchor 4 ('a few natural terms missing') undersells the paraphrase-level coverage here.

5 / 5

Distinctiveness Conflict Risk

The niche is well-anchored ("post-training research for LLMs using the Tinker API"), but the when-clause sweeps in broad phrases like "do research", "choose a model", and "set up training" that could overlap with general ML-training or experiment-tracking skills. Mostly distinct with minor overlap risk against closely related training skills — the anchor-4 fit; not 5 because several trigger terms are not unique to this niche.

4 / 5

Total

19

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (601 lines); consider splitting into references/ and linking

Warning

Total

15

/

16

Passed

Repository
thinking-machines-lab/tinker-cookbook
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.