Fast structured generation and serving for LLMs with RadixAttention prefix caching. Use for JSON/regex outputs, constrained decoding, agentic workflows with tool calls, or when you need 5× faster inference than vLLM with prefix sharing. Powers 300,000+ GPUs at xAI, AMD, NVIDIA, and LinkedIn.
68
82%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
View guide
Critical
Do not install without reviewing
Security
1 critical severity finding. Installing this skill is not recommended: please review these findings carefully if you do intend to do so.
Detected a suspicious URL in the skill instructions that could lead the agent to download and execute malicious scripts or binaries. This includes links to executables from untrusted sources, typosquatting of official packages, URL shorteners that obscure the destination, and personal file hosting services.
This set of URLs includes a third-party pip index (https://flashinfer.ai/whl/cu121/torch2.4/) used to install binary wheels outside of PyPI, which is a higher-risk download source because it can host arbitrary binaries not vetted by official package repositories.
Low
Low-risk findings.
2 low severity findings. Worth noting, but not necessarily harmful.
The skill exposes the agent to untrusted, user-generated content from public third-party sources, creating a risk of indirect prompt injection. This includes browsing arbitrary URLs, reading social media posts or forum comments, and analyzing content from unknown websites.
The required runtime workflow is serving user-provided chat messages/prompts via SGLang (e.g., the OpenAI-compatible `/v1/chat/completions` example in SKILL.md), which means arbitrary free text from outside the operating user can enter the model context as part of `messages`/prompt construction.
The skill fetches instructions or code from an external URL at runtime, and the fetched content directly controls the agent’s prompts or executes code. This dynamic dependency allows the external source to modify the agent’s behavior without any changes to the skill itself.
The skill includes installation steps that fetch and execute remote code (git clone https://github.com/sgl-project/sglang.git and pip install using the external wheel index https://flashinfer.ai/whl/cu121/torch2.4/), which are required to run the skill and thus present a runtime external code-execution dependency.
6da7f7c
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.