CtrlK
BlogDocsLog inGet started
Tessl Logo

speculative-decoding

Accelerate LLM inference using speculative decoding, Medusa multiple heads, and lookahead decoding techniques. Use when optimizing inference speed (1.5-3.6× speedup), reducing latency for real-time applications, or deploying models with limited compute. Covers draft models, tree-based attention, Jacobi iteration, parallel token generation, and production deployment strategies.

60

Quality

70%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./backend/cli/skills/ml-inference/speculative-decoding/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

57%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is well-structured with strong, mostly executable code examples for all three techniques, and the bundle files are genuine and clearly linked. Its main weakness is redundancy: core-concept and advanced sections re-explain material that already lives in the reference files, and there are no validation/benchmark checkpoints around production deployment.

Suggestions

Collapse the Core Concepts and Advanced Patterns sections that duplicate references/medusa.md and references/lookahead.md into one-line summaries with pointers (e.g., 'Medusa tree attention: See references/medusa.md'), keeping only the quick-start code inline.

Replace the pseudocode blocks (speculative_decode, the LookaheadDecoding class, select_draft_model, and the 'Choose the Right Method' if-statements) with an executable decision table, and add the missing imports (torch, AutoTokenizer) so the quick-start examples run as-is.

Add a validation checkpoint to Production Deployment: benchmark tokens/sec before and after enabling speculative decoding and verify output quality is unchanged before shipping.

DimensionReasoningScore

Conciseness

Mostly efficient code-focused content, but the 'When to Use' bullets restate the frontmatter description, the Core Concepts sections re-explain speculative-decoding theory at length, and the Medusa architecture diagram is duplicated from references/medusa.md. Best Practices §1 ('if deploying_new_model: use_method = ...') is non-executable filler.

3 / 5

Actionability

Three near copy-paste-ready quick-start examples plus concrete install commands cover the common cases, with minor gaps: 'torch.float16' is used without importing torch, AutoTokenizer is unimported in the Medusa example, and several blocks (speculative_decode, the LookaheadDecoding class, select_draft_model) are illustrative pseudocode.

4 / 5

Workflow Clarity

The install → quick start → advanced → production progression is present and the method-comparison table plus selection guidance give a rough decision sequence, but there are no validation checkpoints — nothing on benchmarking to confirm the claimed speedup or verifying output quality is unchanged before deploying.

3 / 5

Progressive Disclosure

The two referenced files (references/medusa.md, references/lookahead.md) are real, one level deep, and clearly signaled in 'See Also', but roughly 200 lines of architecture, training, and algorithm detail inlined in Core Concepts and Advanced Patterns duplicate that reference content instead of being delegated to it.

3 / 5

Total

13

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: it explicitly states what the skill does and when to use it, with concrete named techniques and quantified speedup figures. Weaknesses are minor — a noun-heavy technique list rather than distinct actions, and a few missing natural synonyms such as 'throughput' or 'faster generation'.

DimensionReasoningScore

Specificity

The description names the domain ('Accelerate LLM inference') and comprehensively enumerates concrete techniques ('speculative decoding, Medusa multiple heads, and lookahead decoding', 'draft models, tree-based attention, Jacobi iteration'), but offers few distinct action verbs — essentially 'accelerate' and 'covers' — placing it just below the multiple-concrete-actions anchor.

4 / 5

Completeness

Clearly answers both questions: the 'what' ('Accelerate LLM inference using speculative decoding, Medusa multiple heads, and lookahead decoding techniques') and an explicit 'when' with three concrete trigger phrases ('Use when optimizing inference speed..., reducing latency..., or deploying models with limited compute').

5 / 5

Trigger Term Quality

Includes natural phrases users would say ('optimizing inference speed (1.5-3.6x speedup)', 'reducing latency for real-time applications', 'deploying models with limited compute') alongside distinctive technique names, but misses common synonyms like 'speed up generation', 'inference throughput', or 'faster decoding'.

4 / 5

Distinctiveness Conflict Risk

Technique names (Medusa, lookahead decoding, draft models) carve out a clear niche with minimal conflict risk, but the generic trigger 'optimizing inference speed' could also fire for quantization or serving-optimization skills — minor overlap risk.

4 / 5

Total

17

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

14

/

16

Passed

Repository
synthetic-sciences/openscience
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.