CtrlK
BlogDocsLog inGet started
Tessl Logo

inkling

Sample, evaluate, and post-train Inkling and Inkling-Small, Thinking Machines Lab's models built for Tinker. Use this skill whenever the user mentions Inkling, `thinkingmachines/Inkling`, tml-renderers, `tml_v0` / `TmlV0Renderer`, or thinking/reasoning effort — and whenever they are choosing a model, building training data, running evals, setting up SFT or RL, handling parse errors, or working with audio or image inputs for an Inkling model. Inkling has requirements that differ from other Tinker models (mandatory effort conditioning, its own renderer and tokenizer, a learning rate you calibrate yourself), so load this skill before writing any Inkling code.

74

Quality

93%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

82%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A strong, executable skill body: actionable code, clear sequencing with parse-error feedback loops, and well-signaled external references. It earns a top actionability score and near-top marks elsewhere, with only minor trimming and reference-split opportunities. No critical gaps.

Suggestions

Trim a few justificatory asides (e.g. 'which is a reasonable default but should be a deliberate choice', 'deliberately refuses remote URLs') to push conciseness toward the fully-lean anchor.

Consider moving the large pitfalls table and/or effort-preset table into a reference file so SKILL.md reads as a tighter overview pointing one level deep.

Make the post-training checkpoints (e.g. 'watch entropy', 'track all-fail/all-success groups') into an explicit validate-and-act sequence to reach the top workflow-clarity anchor.

DimensionReasoningScore

Conciseness

The body is largely lean and assumes competence — it skips generic explanations of what a renderer or learning rate is — but a few passages add mild justification that could be trimmed (e.g. 'which is a reasonable default but should be a deliberate choice', 'deliberately refuses remote URLs'), so it sits just below the fully-lean anchor.

4 / 5

Actionability

Guidance is concrete and copy-paste ready: executable renderer/sampling/training code blocks, exact effort scalars in a table, runnable script invocations, and a symptoms/causes/fixes pitfalls table covering the common cases.

5 / 5

Workflow Clarity

Multi-step flows (setup → thinking-effort → sampling → training data → post-training → evaluation) are clearly sequenced with explicit error-handling feedback loops for parse errors and explicit 'validate by checking termination/max_tokens first' guidance, but a couple of post-training checkpoints are advisory rather than enforced, leaving a minor gap below the top anchor.

4 / 5

Progressive Disclosure

Structure is well-organized into clear sections with a consolidated Reference list of one-level-deep external links, and no bundle files exist to push detail into; however some reference-style detail (effort presets, pitfalls) is inlined in SKILL.md rather than split out, so it is just shy of the ideal overview-only anchor.

4 / 5

Total

17

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An excellent description: specific, trigger-rich, and explicit about both capabilities and activation conditions, with a clear model-specific niche that minimizes overlap. Third-person voice is maintained throughout. No changes needed.

DimensionReasoningScore

Specificity

The description enumerates many concrete actions — 'Sample, evaluate, and post-train', 'building training data, running evals, setting up SFT or RL, handling parse errors, or working with audio or image inputs' — giving comprehensive, specific coverage of what the skill does.

5 / 5

Completeness

It explicitly answers both 'what' (sample/evaluate/post-train Inkling models) and 'when' (a detailed 'Use this skill whenever...' clause listing model names, renderers, and task categories), matching the top anchor that requires concrete trigger phrases.

5 / 5

Trigger Term Quality

It surfaces natural trigger tokens users actually say — 'Inkling', 'thinkingmachines/Inkling', 'tml-renderers', 'tml_v0 / TmlV0Renderer', 'thinking/reasoning effort' — plus behavioral triggers like 'choosing a model' or 'setting up SFT or RL', covering synonyms and identifiers comprehensively.

5 / 5

Distinctiveness Conflict Risk

The niche is sharply defined by model-specific identifiers (Inkling, tml-renderers, TmlV0Renderer) and a stated rationale for why it differs from other Tinker models (mandatory effort conditioning, own renderer/tokenizer, self-calibrated LR), yielding minimal conflict risk with other skills.

5 / 5

Total

20

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
thinking-machines-lab/tinker-cookbook
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.