CtrlK
BlogDocsLog inGet started
Tessl Logo

vllm-sr-agent-operations

Install, configure, verify, and improve vLLM Semantic Router through its CLI and Router API. Use for deployment, recipe tuning, and single-model/MoM evaluation, with Dashboard verification when requested.

SKILL.md
Quality
Evals
Security

vLLM Semantic Router operations

Work against the user's selected stack and objective. Inspect the installed CLI, running configuration and available backends before choosing an approach. Preserve unrelated workloads, credentials and the user's existing authorization.

Discover the contract

Use vllm-sr --help, command-specific help and vllm-sr config schema. For a running Router, GET /api/v1 advertises its operations and schemas. Management origin, inference listener and public model entrypoint are separate; discover them instead of assuming default ports or a recipe name.

Read only the reference needed for the task:

TaskReference
Install, select hardware/runtime, isolate a stack, open Dashboard accessDeployment
Change live config, activate a recipe, recover a revisionConfiguration
Verify routing, tools, context boundaries or deliveryRoute verification
Improve signal, decision or model-selection policyRecipe tuning
Compare single models and MoM; run a measured optimization loopsr-bench

For installation or an authorized upgrade, use the published stable package unless the user selects another version or the dev channel:

curl -fsSL https://vllm-sr.ai/install.sh | \
  bash -s -- --channel stable --mode cli --runtime skip --no-launch
export PATH="$HOME/.local/bin:$PATH"
vllm-sr --version

Work loop

  • Establish the intended behavior and a small reproducible baseline.
  • For a new stack, initialize and validate config before serve. For an existing stack, derive changes from fresh config get, then validate, plan and apply. Respect restart-required changes and verify the active revision afterward.
  • Preview checks routing without generating an answer. Probe or live evaluation checks actual delivery. Verify the behavior affected by the change, including final output; readiness or HTTP 200 alone is insufficient.
  • Compare the same workload before and after a coherent change. Use sr-bench when capability, cost or latency is the objective. Preserve unsuccessful attempts and distinguish small-sample evidence from a quality claim.

For Dashboard work, exercise the corresponding user flow against the same stack and inspect the resulting artifacts. Leave the user with the active config or recipe, access details, evidence and material limitations. Keep secret values and private request content out of public artifacts.

Repository
vllm-project/semantic-router
Last updated
First committed

Also appears in

vllm-project/semantic-router
In sync

since Sep 22, 2026

Renamed to: vllm-sr

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.