CtrlK
BlogDocsLog inGet started
Tessl Logo

debug

Diagnose training issues with Tinker — slow steps, hanging sessions, output mismatches, error messages, renderer problems, and deployment issues. Use this skill whenever a user reports that training is slow, steps take too long, sessions are hanging, model outputs differ between Tinker and external engines (vLLM, SGLang), they get a confusing error message, training quality is poor (high KL, bad outputs), or they suspect something is wrong. Also trigger when users ask "is this a Tinker issue or my issue?", "is Tinker down?", report unexpected wait times, see output quality regressions, get opaque errors, or want to profile/debug their training or deployment pipeline. This skill walks through systematic triage to determine root cause.

71

Quality

89%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

81%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is an exceptionally actionable, well-sequenced diagnostic playbook with executable code, explicit routing between triage steps, and genuine one-level-deep reference files. Its main weakness is redundancy — duplicated code blocks, a repeated decision-tree branch, and a resolutions table that restates earlier content — which inflates token cost without adding information.

Suggestions

Delete the verbatim duplicate of the "Output correctness" decision-tree branch in the decision tree summary (the five lines from "Using custom merge script?" through "still wrong → Engine-specific issue → Escalate" appear twice in a row).

Drop the repeated pyinstrument install/usage block in Step 4 Option B (it is already given verbatim in Step 3) and reference it instead, e.g. "Run pyinstrument as shown in Step 3".

Trim the "Common resolutions" table to only rows not already covered by the error-message decoder tables and the performance triage steps, or fold the decoder tables into `references/error-reference.md` and keep a single symptom-lookup table inline.

DimensionReasoningScore

Conciseness

The body is mostly dense, useful diagnostic material, but contains systematic duplication: the "Output correctness" decision-tree branch is repeated verbatim twice, the pyinstrument install/usage block appears in both Step 3 and Step 4 Option B, and the "Common resolutions" table restates rows already covered by the error-decoder tables and triage sections. This fits anchor 3 (mostly efficient but could be tightened) rather than 4, where over-explanation is only minor.

3 / 5

Actionability

Fully executable, copy-paste-ready code throughout — environment version check, timing wrappers, GIL monitor thread, token-comparison script, service smoke test, merge/PEFT commands — plus specific version pins and error tables with concrete fixes, directly matching the anchor-5 example.

5 / 5

Workflow Clarity

"Work through these steps in order" establishes an explicit sequence with routing between steps ("High submit time → ... Go to Step 4", "Fast API round-trip → Server-side → Escalate"), validation checkpoints (token-match assert, smoke-test interpretation), per-path decision trees, and escalation checklists listing required artifacts — the feedback-loop structure of the anchor-5 example is present.

5 / 5

Progressive Disclosure

Five real, substantive reference files are one level deep, clearly signaled at point of use ("read `references/async-task-dump.md`") and indexed at the end, with core triage inline and deep detail split out. Falls short of 5 because the inline error-decoder and renderer tables partially duplicate reference material and the duplicated decision-tree block is an organization defect.

4 / 5

Total

17

/

20

Passed

Description

95%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: it names the domain and capabilities concretely, provides exhaustive natural-language trigger phrases, and explicitly covers both what and when in third-person voice. The only minor weakness is that capabilities are conveyed more through scenario listing than through a breadth of distinct diagnostic actions.

DimensionReasoningScore

Specificity

Enumerates several concrete problem categories ("slow steps, hanging sessions, output mismatches, error messages, renderer problems, and deployment issues") and a concrete action ("systematic triage to determine root cause"), but the description leans on scenario enumeration more than a comprehensive list of concrete actions, matching anchor 4 rather than 5.

4 / 5

Completeness

Explicitly answers both what ("Diagnose training issues with Tinker ... walks through systematic triage to determine root cause") and when ("Use this skill whenever a user reports that training is slow...") with concrete trigger phrases, a direct match to the anchor-5 example pattern.

5 / 5

Trigger Term Quality

Covers the natural phrases users would actually say, including synonyms and paraphrases: "training is slow", "steps take too long", "sessions are hanging", "is Tinker down?", "is this a Tinker issue or my issue?", "high KL", "confusing error message" — comprehensive natural-term coverage.

5 / 5

Distinctiveness Conflict Risk

Clear niche — Tinker training/deployment debugging — anchored by product-specific vocabulary (Tinker, vLLM/SGLang mismatch, session heartbeats, high KL) with minimal overlap risk against adjacent debugging skills.

5 / 5

Total

19

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (656 lines); consider splitting into references/ and linking

Warning

Total

15

/

16

Passed

Repository
thinking-machines-lab/tinker-cookbook
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.