Configure NVIDIA NIM platform on OpenShift AI for optimized model inference. Use when: - "Set up NIM on my cluster" - "Configure NGC credentials for NIM" - "I want to deploy a NIM model but haven't set up the platform" - "Create the NIM Account CR" One-time prerequisite before deploying models with NVIDIA NIM runtime via /model-deploy. NOT for deploying models (use /model-deploy instead). NOT for vLLM or Caikit deployments (NIM-specific only).
66
81%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
View guide
Passed
No findings from the security scan
Configure the NVIDIA NIM platform on OpenShift AI. This is a one-time setup that creates NGC credentials and the NIM Account custom resource, enabling NIM-based model deployments via /model-deploy.
Required MCP Server: openshift (OpenShift MCP Server)
Required MCP Tools (from openshift):
resources_get - Check operator installations and existing resourcesresources_list - List resources in a namespaceresources_create_or_update - Create secrets, Account CR, ConfigMapevents_list - Check events for errors during setupOptional MCP Server: rhoai (RHOAI MCP Server)
Optional MCP Tools (from rhoai):
list_data_science_projects - Validate namespace is an RHOAI Data Science Projectlist_serving_runtimes - Verify NIM ServingRuntimes after setupOptional MCP Server: ai-observability (for get_gpu_info to verify GPU availability)
Common prerequisites (KUBECONFIG, OpenShift+RHOAI cluster, verification protocol): See skill-conventions.md.
Required User Input:
Additional cluster requirements:
Use this skill when you need to:
Do NOT use this skill when:
/model-deploy after NIM setup is complete)/model-deploy directly)/serving-runtime-config)If the rhoai MCP server is available, validate that the target namespace is an RHOAI Data Science Project:
MCP Tool: list_data_science_projects (from rhoai)
If the namespace is not in the project list, warn: "Namespace [namespace] is not a Data Science Project. NIM setup may not work correctly. Consider creating a Data Science Project first."
If rhoai MCP is not available, skip this check and proceed.
Document Consultation (read before verifying operators):
Check that the NVIDIA GPU Operator and NFD Operator are installed and healthy.
MCP Tool: resources_get (from openshift)
Parameters:
apiVersion: "operators.coreos.com/v1alpha1" - REQUIREDkind: "ClusterServiceVersion" - REQUIREDnamespace: "nvidia-gpu-operator" - REQUIRED (namespace where GPU Operator CSV is installed)name: the CSV name matching "gpu-operator-certified" prefixExpected Output: ClusterServiceVersion object with status.phase: "Succeeded"
Repeat for NFD Operator:
namespace: "openshift-nfd"name: the CSV name matching "nfd" prefixError Handling:
status.phase != "Succeeded" -> Report current phase and suggest troubleshootingAsk the user for their NGC API key. This key is used for two purposes:
nvcr.io (image pull secret)Ask the user:
To set up NIM, I need your NVIDIA NGC API key.
You can generate one at: https://ngc.nvidia.com/setup/api-key
Please provide:
1. Your NGC API key
2. The target namespace for NIM resources (e.g., "my-ai-project")WAIT for user to provide the NGC API key and namespace.
SECURITY: Store the key in memory only for the duration of this skill. Never echo or display the actual key value in output.
Generate and display the docker-registry Secret YAML for pulling NIM images from nvcr.io.
Show the user the Secret manifest (with API key value redacted):
apiVersion: v1
kind: Secret
metadata:
name: ngc-image-pull-secret
namespace: [namespace]
type: kubernetes.io/dockerconfigjson
data:
.dockerconfigjson: [base64-encoded docker config for nvcr.io]Note: The .dockerconfigjson contains:
nvcr.io$oauthtoken[NGC API key - REDACTED in display]Ask: "Should I create this image pull secret in namespace [namespace]? (yes/no)"
WAIT for explicit user confirmation.
MCP Tool: resources_create_or_update (from openshift)
Parameters:
manifest: full Secret manifest as JSON string - REQUIRED
namespace: user-specified namespace - REQUIRED
"my-ai-project"Expected Output: Created Secret object with metadata.uid
Error Handling:
ngc-image-pull-secret already exists. Should I update it? (yes/no)"Generate and display the generic Secret YAML for the NGC API key used at runtime.
Show the user the Secret manifest (with API key value redacted):
apiVersion: v1
kind: Secret
metadata:
name: ngc-api-key
namespace: [namespace]
type: Opaque
stringData:
NGC_API_KEY: "[REDACTED]"Ask: "Should I create this API key secret in namespace [namespace]? (yes/no)"
WAIT for explicit user confirmation.
MCP Tool: resources_create_or_update (from openshift)
Parameters:
manifest: full Secret manifest as JSON string - REQUIREDnamespace: user-specified namespace - REQUIREDExpected Output: Created Secret object with metadata.uid
Error Handling:
Generate and display the NIM Account custom resource that manages the NIM platform lifecycle.
Show the user the Account CR manifest:
apiVersion: nim.opendatahub.io/v1
kind: Account
metadata:
name: nim-account
namespace: [namespace]
spec:
apiKeySecret:
name: ngc-api-key
imagePullSecret:
name: ngc-image-pull-secretAsk: "Should I create this NIM Account CR in namespace [namespace]? (yes/no)"
WAIT for explicit user confirmation.
MCP Tool: resources_create_or_update (from openshift)
Parameters:
manifest: full Account CR manifest as JSON string - REQUIREDnamespace: user-specified namespace - REQUIREDExpected Output: Created Account object with metadata.uid
Error Handling:
nim.opendatahub.io/v1 Account) -> Report: "NIM CRD not available. Ensure Red Hat OpenShift AI operator is installed and includes NIM support."Ask: "Would you like to customize which NIM models appear in the catalog? (yes/no, default: no)"
If user says no -> Skip to Step 7 (default catalog is used).
If user says yes:
Show the user the ConfigMap template:
apiVersion: v1
kind: ConfigMap
metadata:
name: nim-model-catalog
namespace: [namespace]
data:
model-catalog.json: |
[
{
"name": "[model-name]",
"displayName": "[display-name]",
"shortDescription": "[description]"
}
]Ask user which models to include, generate the ConfigMap, and confirm before creating.
MCP Tool: resources_create_or_update (from openshift)
Check that the NIM platform is ready for model deployments.
Step 7a: Check Account CR Status
MCP Tool: resources_get (from openshift)
Parameters:
apiVersion: "nim.opendatahub.io/v1" - REQUIREDkind: "Account" - REQUIREDnamespace: user-specified namespace - REQUIREDname: "nim-account" - REQUIREDExpected Output: Account object with status.conditions showing ready state
Step 7b: Verify NIM ServingRuntimes
MCP Tool: list_serving_runtimes (from rhoai) - preferred if rhoai MCP available
Parameters:
namespace: user-specified namespace - REQUIREDinclude_templates: falseFallback MCP Tool: resources_list (from openshift)
apiVersion: "serving.kserve.io/v1alpha1", kind: "ServingRuntime", namespace: user-specified namespaceExpected Output: List of ServingRuntime objects including NIM runtimes
Step 7c: (Optional) GPU Inventory Check
If ai-observability MCP server is available, use get_gpu_info to report cluster GPU inventory.
Report results showing: Account CR status, credentials status (created/existing), available NIM ServingRuntimes, GPU inventory (if available), and next steps (/model-deploy).
On failure: Report Account CR status details and error message. Suggest troubleshooting steps: check Account CR events, verify NGC API key validity, check OpenShift AI operator logs. Ask if user wants help troubleshooting.
For common issues (GPU scheduling, OOMKilled, image pull errors, RBAC), see common-issues.md.
Error: Account CR status.conditions shows pending state indefinitely
Cause: NGC credentials are invalid, expired, or the RHOAI operator cannot reach NGC services.
Solution:
events_list filtered by namespace to find events related to the Account resource/nim-setup with new credentialsError: ClusterServiceVersion for gpu-operator-certified not found
Cause: NVIDIA GPU Operator was not installed from OperatorHub.
Solution:
Succeeded phasenvidia.com/gpu resources on nodes/nim-setupError: resources_list for ServingRuntimes returns no NIM runtimes
Cause: Account CR is not yet ready, or the RHOAI operator version does not include NIM support.
Solution:
See Prerequisites for the complete list of required and optional MCP tools.
/model-deploy - Deploy a model using NIM runtime after setup is complete/serving-runtime-config - Configure custom serving runtimes if NIM doesn't fitWhen handing off to /model-deploy after NIM setup, note these NIM-specific considerations:
nim:// URI scheme which the deploy_model RHOAI tool may not recognize. If this happens, /model-deploy will fall back to creating the InferenceService via OpenShift direct with the NIM container image, NGC credentials, and NIM-specific env vars.latest NIM image tag may bundle a CUDA version incompatible with the cluster's GPU drivers. Always recommend a specific tag (e.g., 1.8.3) matched to the GPU driver version.NIM_MAX_MODEL_LEN=16384 for T4/A10 GPUs./model-deploy will automatically detect and add tolerations after deployment.See skill-conventions.md for general HITL and security conventions.
Skill-specific checkpoints:
See nim-setup examples for a complete first-time NIM setup walkthrough.
e46c4fa
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.