CtrlK
BlogDocsLog inGet started
Tessl Logo

serving-llms-vllm

Serves LLMs with high throughput using vLLM's PagedAttention and continuous batching. Use when deploying production LLM APIs, optimizing inference latency/throughput, or serving models with limited GPU memory. Supports OpenAI-compatible endpoints, quantization (GPTQ/AWQ/FP8), and tensor parallelism.

64

Quality

76%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./backend/cli/skills/ml-inference/vllm/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

61%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, largely actionable skill with real reference files and mostly executable commands, but held back by comment-only placeholder steps, some padding that explains things Claude already knows, and a batch-inference workflow lacking any output validation or feedback loop.

Suggestions

Add a validation/feedback step to Workflow 2 (batch inference): after Step 4, check for empty or truncated outputs (e.g., verify token counts > 0 and no stop-string truncation) and retry failed prompts, which would lift workflow_clarity.

Replace comment-only placeholder blocks (load-test setup in Workflow 1 Step 2, model search and accuracy verification in Workflow 3) with runnable code, or move them into the referenced troubleshooting/quantization files.

Trim known-concept explanations such as 'vLLM automatically batches requests for efficiency / No need to manually chunk prompts' and the generic checklist preamble to reduce token spend.

DimensionReasoningScore

Conciseness

Mostly efficient code/commands, but includes unnecessary padding: comment-only placeholder blocks ("# Create test_load.py with sample requests", workflow 3's "# Compare quantized vs non-quantized responses"), and explanations of concepts Claude already knows ("vLLM automatically batches requests for efficiency... No need to manually chunk prompts"). Not a 4 because several sections need tightening or removal rather than just one or two spots.

3 / 5

Actionability

Mostly executable: concrete `vllm serve` commands with flags, runnable Python for offline inference and batch processing, docker commands, specific metric names and thresholds. Not a 5 because a few steps (load-test setup, quantized-model search, accuracy verification) are comment-only placeholders rather than copy-paste-ready code.

4 / 5

Workflow Clarity

Workflows 1 and 3 have clear numbered steps and workflow 1 includes explicit verification checkpoints ("Verify TTFT < 500ms and throughput > 100 req/sec", "No OOM errors in logs"), but workflow 2 is a batch operation whose final step just processes results with no validation or feedback loop — the rubric caps batch operations lacking validation at 3.

3 / 5

Progressive Disclosure

Good structure: four clearly signaled one-level-deep references under "Advanced topics", all of which exist as real files in references/. Not a 5 because the 378-line body inlines extensive workflow code that could partly live in the reference files, leaving the main file less of an overview.

4 / 5

Total

14

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: concise, third-person, specific about capabilities, with an explicit 'Use when' clause containing multiple concrete trigger phrases. The only minor gap is a few missing natural synonyms (e.g., 'hosting', 'inference server').

DimensionReasoningScore

Specificity

Lists multiple specific concrete capabilities — "Serves LLMs with high throughput using vLLM's PagedAttention and continuous batching", "OpenAI-compatible endpoints, quantization (GPTQ/AWQ/FP8), and tensor parallelism" — in third person, with comprehensive coverage of the skill's scope. Not a 4 because no meaningful capability gaps exist for this domain.

5 / 5

Completeness

Explicitly answers both: what ("Serves LLMs with high throughput... Supports OpenAI-compatible endpoints, quantization, and tensor parallelism") and when ("Use when deploying production LLM APIs, optimizing inference latency/throughput, or serving models with limited GPU memory") with concrete trigger phrases. Matches the anchor-5 exemplar structure.

5 / 5

Trigger Term Quality

Good natural-term coverage: "deploying production LLM APIs", "optimizing inference latency/throughput", "serving models", "limited GPU memory", plus the vLLM name itself. Not a 5 because common variations like "hosting", "inference server", or "LLM API" phrasings are absent.

4 / 5

Distinctiveness Conflict Risk

Clear niche anchored to a specific tool (vLLM) with tool-specific techniques (PagedAttention, tensor parallelism), so it is unlikely to fire for unrelated skills. Not a 4 because the tool name makes even overlap with adjacent serving skills (TGI, TensorRT-LLM) minimal.

5 / 5

Total

19

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

14

/

16

Passed

Repository
synthetic-sciences/openscience
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.