CtrlK
BlogDocsLog inGet started
Tessl Logo

serving-llms-vllm

Serves LLMs with high throughput using vLLM's PagedAttention and continuous batching. Use when deploying production LLM APIs, optimizing inference latency/throughput, or serving models with limited GPU memory. Supports OpenAI-compatible endpoints, quantization (GPTQ/AWQ/FP8), and tensor parallelism.

65

Quality

78%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Critical

Do not install without reviewing

Fix and improve this skill with Tessl

tessl review fix ./12-inference-serving/vllm/SKILL.md
SKILL.md
Quality
Evals
Security

Security

1 critical severity finding. Installing this skill is not recommended: please review these findings carefully if you do intend to do so.

Critical

E006: Malicious code pattern detected in skill scripts.

What this means

Detected high-risk code patterns in the skill content — including its prompts, tool definitions, and resources — such as data exfiltration, backdoors, remote code execution, credential theft, system compromise, supply chain attacks, and obfuscation techniques.

Why it was flagged

The docs repeatedly instruct users to enable --trust-remote-code (and show it in production examples), which permits executing arbitrary code from remote model repositories and therefore introduces a high risk of remote code execution/backdoor abuse; no explicit data exfiltration code was found in the docs.

Report incorrect finding

Low

Low-risk findings.

2 low severity findings. Worth noting, but not necessarily harmful.

Low

W011: Third-party content exposure detected (indirect prompt injection risk).

What this means

The skill exposes the agent to untrusted, user-generated content from public third-party sources, creating a risk of indirect prompt injection. This includes browsing arbitrary URLs, reading social media posts or forum comments, and analyzing content from unknown websites.

Why it was flagged

This skill’s workflow starts a vLLM OpenAI-compatible server that ingests runtime `messages[].content` from external callers into the model prompt context (e.g., any user request text, which can include free-form outsider-provided prompt injection), so outsider free text can reach the LLM via inference inputs.

Low

W012: Unverifiable external dependency detected (runtime URL that controls agent).

What this means

The skill fetches instructions or code from an external URL at runtime, and the fetched content directly controls the agent’s prompts or executes code. This dynamic dependency allows the external source to modify the agent’s behavior without any changes to the skill itself.

Why it was flagged

The skill shows running vllm with remote model identifiers (e.g., meta-llama/Llama-3-8B-Instruct and TheBloke/Llama-2-70B-AWQ) which vLLM will fetch at runtime and the docs explicitly mention using --trust-remote-code to allow executing code from those remote model repositories, creating a high-confidence remote-code execution risk.

Repository
Orchestra-Research/AI-Research-SKILLs
Audited
Security analysis
Snyk

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.