CtrlK
BlogDocsLog inGet started
Tessl Logo

knowledge-distillation

Compress large language models using knowledge distillation from teacher to student models. Use when deploying smaller models with retained performance, transferring GPT-4 capabilities to open-source models, or reducing inference costs. Covers temperature scaling, soft targets, reverse KLD, logit distillation, and MiniLLM training strategies.

59

Quality

69%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./backend/cli/skills/ml-training/knowledge-distillation/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

50%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is rich in concrete, mostly executable code and covers the key distillation strategies, but it is significantly overlong: training-loop boilerplate is repeated, basics Claude already knows are re-explained, and the MiniLLM reference file is duplicated inline rather than linked. It also lacks validation checkpoints for the batch training workflows it describes.

Suggestions

Replace the inline "MiniLLM (Reverse KLD)" section with a one-line pointer to references/minillm.md (e.g., "**MiniLLM / reverse KLD**: See [references/minillm.md](references/minillm.md)"), and cut the Core Concepts temperature-scaling and forward-vs-reverse-KL explanations Claude already knows.

Make every code snippet executable: replace the invalid "train_data = {\"teacher_generated\": 70%}" dict and the pseudocode "student = distill(teacher, student, epochs=5)" with runnable equivalents, and define or remove "calculate_similarity" in the Evaluation section.

Add validation checkpoints to the training workflow, e.g., "Verify the combined loss decreases over the first ~500 steps before continuing" and "Evaluate the student against the teacher on held-out prompts and only save/deploy when quality is acceptable".

Consolidate the three near-identical training loops (Quick Start, Strategy 1, Production Deployment) into one canonical example to remove repeated boilerplate.

DimensionReasoningScore

Conciseness

The ~445-line body repeats the same teacher-forward/student-forward/loss/backward loop in at least three places and includes a "Core Concepts" section explaining temperature scaling with hand-computed softmax values — concepts Claude already knows. This matches "noticeably verbose; several unnecessary explanations or padded sections" rather than a 3, where padding would be incidental rather than pervasive.

2 / 5

Actionability

Most guidance is concrete, executable PyTorch/transformers code (the distillation loss, the DistillationTrainer subclass, multi-teacher averaging). It is not a 5 because several snippets are non-executable: "train_data = {\"teacher_generated\": 70%}" is invalid Python, Strategy 2's "student = distill(teacher, student, epochs=5)" calls undefined functions, and the Evaluation section uses an undefined "calculate_similarity".

4 / 5

Workflow Clarity

Content is presented as parallel code recipes rather than a sequenced workflow, and there are no validation checkpoints or feedback loops (e.g., verify distillation loss is decreasing, evaluate student against teacher before deploying) for what is a long batch training operation. Per the guideline capping batch operations without validation at 3, this cannot score 4 despite the recipes themselves being individually clear.

3 / 5

Progressive Disclosure

Section headers give the body reasonable structure, but the 334-line "references/minillm.md" bundle file is never linked or mentioned from the body, and the body's own "MiniLLM (Reverse KLD)" section duplicates that file's material inline. This matches "references present but not clearly signaled; content that should be separate is inline" — structure exists (not a 2), but the only bundle file is invisible from SKILL.md (not a 4).

3 / 5

Total

12

/

20

Passed

Description

88%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: it clearly states what the skill does and gives an explicit, concrete "Use when" clause with multiple triggering scenarios plus a comprehensive technique list. The only weaknesses are a few missing natural synonyms (e.g., "model compression") and mild overlap with inference-cost-reduction skills like quantization.

DimensionReasoningScore

Specificity

The description states the core action ("Compress large language models using knowledge distillation from teacher to student models") and comprehensively enumerates the technique coverage ("temperature scaling, soft targets, reverse KLD, logit distillation, and MiniLLM training strategies"), matching the anchor for multiple specific concrete actions with comprehensive coverage. It is not a 4 because the technique list leaves no significant gaps in the skill's stated domain.

5 / 5

Completeness

It explicitly answers both questions: the "what" ("Compress large language models using knowledge distillation from teacher to student models") and the "when" via a concrete "Use when..." clause with three specific triggering scenarios. This matches the top anchor; the "when" is explicit and specific, ruling out a 4.

5 / 5

Trigger Term Quality

Natural phrases like "deploying smaller models with retained performance", "transferring GPT-4 capabilities to open-source models", and "reducing inference costs" map well to what a user would actually say. It is not a 5 because common variations such as "model compression", "distill a model", or "teacher-student training" are absent.

4 / 5

Distinctiveness Conflict Risk

Niche-specific terms (reverse KLD, MiniLLM, logit distillation, teacher-to-student) make it largely distinct from adjacent skills, but "reducing inference costs" overlaps with quantization/pruning skills, creating minor conflict risk. Not a 5 due to that overlap; clearly above a 3 given the KD-specific triggers.

4 / 5

Total

18

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

14

/

16

Passed

Repository
synthetic-sciences/openscience
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.