Configure custom ServingRuntime CRs on OpenShift AI for model serving frameworks not covered by built-in runtimes. Use when: - "Create a custom serving runtime" - "I need a runtime for ONNX / Triton / custom framework" - "Customize vLLM runtime parameters" - "What serving runtimes are available?" - "Add a custom container image for model serving" Handles listing existing runtimes, creating new ServingRuntime CRs, and validating compatibility with target models. NOT for deploying models (use /model-deploy after runtime is configured). NOT for NIM platform setup (use /nim-setup).
74
92%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
View guide
Low
Low-risk findings worth noting
Configure custom ServingRuntime custom resources on Red Hat OpenShift AI. Use when built-in runtimes (vLLM, NIM, Caikit+TGIS) do not support the target model framework, or when customizing an existing runtime's parameters (env vars, model format, container image).
Required MCP Server: openshift (OpenShift MCP Server)
Required MCP Tools (from openshift):
resources_get - Inspect existing ServingRuntime CRs in detailresources_list - List ServingRuntime and ClusterServingRuntime CRs (OpenShift fallback)resources_create_or_update - Create fully custom ServingRuntime CR (when not using templates, or as fallback)Preferred MCP Server: rhoai (RHOAI MCP Server) — used when available, automatic OpenShift fallback on failure
Preferred MCP Tools (from rhoai):
list_serving_runtimes - List available runtimes and platform templates with supported model formatscreate_serving_runtime - Instantiate a serving runtime from a platform template (no YAML needed)list_data_science_projects - Validate namespace is an RHOAI projectOptional MCP Server: ai-observability (AI Observability MCP)
Optional MCP Tools (from ai-observability):
list_models - Verify deployed models use the new runtimeCommon prerequisites (KUBECONFIG, OpenShift+RHOAI cluster, KServe, verification protocol): See skill-conventions.md.
Fallback templates: See openshift-fallback-templates.md for OpenShift YAML templates used when RHOAI tools are unavailable.
Use this skill when you need to:
Do NOT use this skill when:
/model-deploy)/nim-setup)/debug-inference)Ask the user for:
MCP Tool: list_data_science_projects (from rhoai)
Parameters: none
Verify the user-specified namespace is an RHOAI Data Science Project.
If rhoai unavailable or returns error: Use resources_list (from openshift) with apiVersion: v1, kind: Namespace, labelSelector: opendatahub.io/dashboard=true.
Error Handling:
[namespace] is not an RHOAI Data Science Project. Use /ds-project-setup to create one, or specify a different namespace." WAIT for user decision.Ask the user for:
Document Consultation (read before listing runtimes):
MCP Tool: list_serving_runtimes (from rhoai)
Parameters:
namespace: validated namespace from Step 1 - REQUIREDinclude_templates: true - REQUIRED (shows both existing runtimes and platform templates)If rhoai unavailable or returns error: Use resources_list (from openshift) with apiVersion: serving.kserve.io/v1alpha1, kind: ServingRuntime, namespace: [namespace] for namespace runtimes, and kind: ClusterServingRuntime for platform templates. Filter by label opendatahub.io/dashboard=true and check spec.supportedModelFormats for compatibility.
Present findings in a table:
| Runtime Name | Model Format | Source | Requires Instantiation |
|---|---|---|---|
| [name] | [format] | namespace / template | [true/false] |
The response distinguishes between:
source: "namespace") - ready to use with /model-deploysource: "template", requires_instantiation: true) - must be instantiated firstIf an existing runtime fits the user's need, recommend using it directly with /model-deploy. If a platform template fits, offer to instantiate it (Step 5 alternative). Otherwise, proceed to Step 3 for custom runtime creation.
WAIT for user to confirm whether to create a new runtime, instantiate a template, or customize an existing one.
Based on the user's framework and model requirements, determine the ServingRuntime spec.
If customizing an existing runtime:
MCP Tool: resources_get (from openshift)
Parameters:
apiVersion: "serving.kserve.io/v1alpha1" - REQUIREDkind: "ServingRuntime" - REQUIREDnamespace: user-specified namespace - REQUIREDname: name of the existing runtime to customize - REQUIREDExtract the current spec as a starting point. Present the current configuration and ask what the user wants to change.
If the user requests a runtime for an unfamiliar framework -> Trigger live doc lookup:
Collect runtime parameters:
| Parameter | Value | Source |
|---|---|---|
| Runtime name | [name] | user input |
| Container image | [image:tag] | user input / doc lookup |
| Model format name | [format] | user input / doc lookup |
| Supported protocol versions | [v1, v2, grpc-v2] | user input / default |
| Multi-model serving | [true/false] | default: false (single-model) |
| Environment variables | [list] | user input |
| GPU resource requirements | [limits] | user input |
WAIT for user to confirm or modify parameters.
Generate the ServingRuntime manifest using values from Steps 2-3.
apiVersion: serving.kserve.io/v1alpha1
kind: ServingRuntime
metadata:
name: [runtime-name]
namespace: [namespace]
labels:
opendatahub.io/dashboard: "true"
annotations:
openshift.io/display-name: "[Display Name]"
spec:
supportedModelFormats:
- name: [model-format-name]
version: "[version]"
autoSelect: true
multiModel: false
containers:
- name: kserve-container
image: [container-image:tag]
ports:
- containerPort: 8080
protocol: TCP
env:
- name: [ENV_VAR_NON_SECRET]
value: "[non-sensitive-value]"
- name: [SECRET_ENV_VAR]
valueFrom:
secretKeyRef:
name: [k8s-secret-name]
key: [secret-key-name]
resources:
limits:
nvidia.com/gpu: "[gpu-count]"
requests:
cpu: "[cpu]"
memory: "[memory]"Display the ServingRuntime YAML to the user, redacting any sensitive values.
Ask: "Proceed with creating this ServingRuntime? (yes/no/modify)"
WAIT for explicit confirmation.
If instantiating from a platform template (user chose a template from Step 2):
MCP Tool: create_serving_runtime (from rhoai)
Parameters:
namespace: target namespace - REQUIREDtemplate_name: name of the template to instantiate (e.g., "vllm-cuda-runtime-template") - REQUIREDThe response includes the created runtime name, display name, and supported model formats.
If rhoai unavailable or returns error: Use resources_get (from openshift) to fetch the ClusterServingRuntime template, copy its spec to a namespace-scoped ServingRuntime, and create via resources_create_or_update (from openshift). See openshift-fallback-templates.md for the pattern.
If creating a fully custom runtime (custom container image, non-template configuration):
MCP Tool: resources_create_or_update (from openshift)
Parameters:
manifest: full ServingRuntime manifest as JSON string - REQUIREDnamespace: user-specified namespace - REQUIREDError Handling:
/ds-project-setup[name] already exists. Update it? (yes/no)"MCP Tool: list_serving_runtimes (from rhoai)
Parameters:
namespace: user-specified namespace - REQUIREDinclude_templates: falseVerify the runtime appears in the namespace runtime list.
If rhoai unavailable or returns error: Use resources_list (from openshift) with apiVersion: serving.kserve.io/v1alpha1, kind: ServingRuntime, namespace: [namespace] for namespace runtimes, and kind: ClusterServingRuntime for platform templates. Filter by label opendatahub.io/dashboard=true and check spec.supportedModelFormats for compatibility.
For detailed inspection:
MCP Tool: resources_get (from openshift)
Parameters:
apiVersion: "serving.kserve.io/v1alpha1" - REQUIREDkind: "ServingRuntime" - REQUIREDnamespace: user-specified namespace - REQUIREDname: the created runtime name - REQUIREDReport results showing: runtime name, namespace, model format, container image, and next steps (/model-deploy to deploy a model using this runtime).
For common issues (GPU scheduling, OOMKilled, image pull errors, RBAC), see common-issues.md.
Error: InferenceService status shows "Unknown" or runtime not matched
Cause: The modelFormat.name in the InferenceService does not match any supportedModelFormats[].name in available ServingRuntimes.
Solution:
opendatahub.io/dashboard: "true" labelError: InferenceService created but health checks fail, endpoint returns connection refused
Cause: The containerPort in the ServingRuntime does not match the port the serving framework actually listens on.
Solution:
containerPort in the ServingRuntime specSee Prerequisites for the complete list of required and optional MCP tools.
/model-deploy - Deploy a model using the configured runtime/nim-setup - NIM platform setup (if NIM runtime is needed instead)/debug-inference - Troubleshoot InferenceService failures after deploymentSee skill-conventions.md for general HITL and security conventions.
Skill-specific checkpoints:
/ds-project-setupe46c4fa
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.