Pick the right inference backend for a MiniCPM5-1B or MiniCPM5-2B checkpoint and route to a backend-specific cookbook skill. Use when the user wants to deploy / serve / chat-with / benchmark a MiniCPM5 model and has not yet committed to a specific engine, or when they say "deploy MiniCPM5", "run MiniCPM5", "serve MiniCPM5", "MiniCPM5 推理", "部署 MiniCPM5".
76
95%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
Passed
No findings from the security scan
You're being asked to deploy / serve / chat-with a MiniCPM5-1B or MiniCPM5-2B checkpoint. Your job is to pick exactly one backend skill below based on the user's hardware, format, and goal, then invoke that skill rather than improvising.
Before picking a backend, you MUST know:
| Variable | Example | Where to ask |
|---|---|---|
MODEL_PATH | HF id openbmb/MiniCPM5-2B or openbmb/MiniCPM5-1B or a local path | "Which checkpoint? HF id or local path?" |
| Hardware | NVIDIA GPU / Apple Silicon / CPU only | infer from context, otherwise ask |
| Goal | "interactive chat" / "OpenAI server" / "Python script" / "benchmark" | infer from context |
| Variant | HF repo | Use with |
|---|---|---|
| HF fp16 (recommended) | openbmb/MiniCPM5-2B or openbmb/MiniCPM5-1B | transformers / vllm (no --quantization) / vllm-ascend / sglang / any minicpm5-finetune-* |
| GGUF F16 / Q8_0 / Q4_K_M | openbmb/MiniCPM5-2B-GGUF or openbmb/MiniCPM5-1B-GGUF | minicpm5-deploy-llama-cpp / -ollama / -lmstudio |
| MLX (Apple Silicon) | openbmb/MiniCPM5-2B-MLX or openbmb/MiniCPM5-1B-MLX | minicpm5-deploy-mlx |
LiteRT-LM .litertlm (Android / iOS / desktop / IoT, CPU + GPU) | litert-community/MiniCPM5-2B or litert-community/MiniCPM5-1B | minicpm5-deploy-litert |
If the user has a local copy, accept any directory path that contains config.json and model.safetensors (or the equivalent GGUF / MLX layout).
| User says / wants | Hardware | Format | → Skill to invoke |
|---|---|---|---|
| "Quick Python script" / "one-shot generation" / "no server" | any GPU or CPU | HF safetensors | minicpm5-deploy-transformers |
| "OpenAI server" / "production serving" / "high QPS" | NVIDIA GPU | HF safetensors | minicpm5-deploy-vllm |
| "vLLM-Ascend" / "Ascend NPU" / "CANN" / "torch_npu" | Huawei Ascend NPU | HF safetensors | minicpm5-deploy-vllm-ascend |
| "RadixAttention" / "prefix cache" / "batched eval" | NVIDIA GPU | HF safetensors | minicpm5-deploy-sglang |
| "GGUF" / "llama.cpp" / "llama-cli" / "CPU only" | any CPU + optional GPU | GGUF | minicpm5-deploy-llama-cpp |
| "Ollama" / "ollama run" / "Modelfile" | macOS / Linux laptop | GGUF | minicpm5-deploy-ollama |
| "LM Studio" / "desktop GUI" | macOS / Windows / Linux | GGUF or MLX | minicpm5-deploy-lmstudio |
| "MLX" / "Apple Silicon native" / "fastest on Mac" | Apple Silicon | MLX | minicpm5-deploy-mlx |
| "Android" / "on-device app" / "Edge Gallery" / "LiteRT" / "LiteRT-LM" / "litertlm" / "Raspberry Pi" | Android phone, iPhone, desktop or IoT board (CPU or GPU) | LiteRT-LM .litertlm | minicpm5-deploy-litert |
If the user has not specified any of the above and asks "how do I run this?":
minicpm5-deploy-vllm.minicpm5-deploy-vllm-ascend.minicpm5-deploy-transformers.minicpm5-deploy-ollama (easiest) or minicpm5-deploy-mlx (fastest).minicpm5-deploy-litert.minicpm5-deploy-llama-cpp (Q4_K_M).Once you've picked a backend skill, invoke that skill with MODEL_PATH set. Do NOT inline the backend's commands here — each backend has its own pitfalls (mandatory flags, env-var pins, install order) that the dedicated skill handles. The user explicitly does NOT want you to "improvise" — read the picked sub-skill in full first.
Whichever backend you pick, after launch run this universal sanity check:
# Replace localhost:PORT with the backend's actual port (default in each sub-skill)
curl http://localhost:PORT/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "MiniCPM5-2B",
"messages": [{"role":"user","content":"1+1=?"}],
"temperature": 1.0, "top_p": 0.95, "max_tokens": 64,
"chat_template_kwargs": {"enable_thinking": true}
}'Expected: HTTP 200 with choices[0].message.content containing "2".
minicpm5-deploy-litert also serves this endpoint: after litert-lm import … minicpm5-2b, litert-lm serve listens on port 9379 (guide); use "model": "minicpm5-2b", add "reasoning_effort": "none" for a direct answer, and cap with max_completion_tokens.
These are common to multiple backends — surface to the user up front:
temperature=1.0, top_p=0.95. MiniCPM5-1B also supports No-think with enable_thinking=false + temperature=0.7, top_p=0.95.max_position_embeddings=131072, rope_theta=5e6, no rope-scaling. Pass --max-model-len 131072 (vLLM) / --context-length 131072 (SGLang) / -c 131072 (llama.cpp) to use the full window. Lower if VRAM is tight.tie_word_embeddings=false. Tools that assume the Llama tied default (e.g. mlx_lm.convert < 0.31) will silently drop lm_head → output collapses to random tokens. The MLX skill bakes in the fix.Each sub-skill is paired with a one-page cookbook in docs/deployment/. The skill is the machine-readable shortcut; the cookbook is the human-readable reference. Both are kept in sync.
316cfb1
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.