Troubleshoot failed or slow InferenceService deployments on OpenShift AI. Use when: - "My InferenceService won't start" - "Model deployment is stuck" - "Inference endpoint returns errors" - "Model is slow / high latency" - "GPU scheduling failed for my model" Progressive diagnosis: status conditions, events, pod logs, GPU health, and observability analysis. NOT for deploying models (use /model-deploy). NOT for creating runtimes (use /serving-runtime-config).
75
94%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
View guide
Low
Low-risk findings worth noting
Troubleshoot failed, stuck, or slow InferenceService deployments on Red Hat OpenShift AI. Performs progressive diagnosis through status conditions, events, pod logs, related resources, and optional observability analysis. Follows a 6-step diagnosis pattern with human-in-the-loop confirmation at each step.
Required MCP Server: openshift (OpenShift MCP Server)
Required MCP Tools (from openshift):
resources_get - Get ServingRuntime, NIM Account CR, InferenceService detailsresources_list - List InferenceServices (OpenShift fallback)pods_list - Find predictor/transformer podspods_log - Retrieve container logsevents_list - Check events for errorsPreferred MCP Server: rhoai (RHOAI MCP Server) — used when available, automatic OpenShift fallback on failure
Preferred MCP Tools (from rhoai):
list_inference_services - List deployed models with structured status dataget_inference_service - Get detailed deployment status (conditions, endpoint, ready state)get_model_endpoint - Quick check if endpoint is available (early diagnostic)Optional MCP Server: ai-observability (AI Observability MCP)
Optional MCP Tools (from ai-observability):
get_deployment_info - Check model initialization statusanalyze_vllm - Analyze vLLM performance bottlenecks (latency, throughput, errors, token rates)chat_vllm - Conversational follow-up on vLLM metrics during diagnosisget_gpu_info - GPU inventory and utilizationanalyze_openshift - Check GPU health with "GPU & Accelerators" categoryquery_tempo_tool - Trace request latency by service/operation/time rangeget_trace_details_tool - Get detailed span-level info for a specific trace IDexecute_promql - Run custom PromQL queries for metrics not covered by standard analysiskorrel8r_get_correlated - Correlate signals (logs, traces, metrics, alerts) across a pod/namespace for root cause analysisCommon prerequisites (KUBECONFIG, OpenShift+RHOAI cluster, KServe, verification protocol): See skill-conventions.md.
Fallback templates: See openshift-fallback-templates.md for OpenShift YAML templates used when RHOAI tools are unavailable.
Additional cluster requirements:
Use this skill when you need to:
Do NOT use this skill when:
/model-deploy)/ai-observability)/serving-runtime-config)/nim-setup)Ask the user:
If user says "list all" or is unsure:
MCP Tool: list_inference_services (from rhoai)
Parameters:
namespace: user-specified namespace - REQUIREDverbosity: "standard" - OPTIONALIf rhoai unavailable or returns error: Use resources_list (from openshift) with apiVersion: serving.kserve.io/v1beta1, kind: InferenceService, namespace: [namespace].
Present InferenceServices with their status:
| Name | Runtime | Ready | URL | Age |
|---|---|---|---|---|
| [name] | [runtime] | [True/False/Unknown] | [url or "N/A"] | [age] |
WAIT for user to select which InferenceService to debug.
MCP Tool: get_inference_service (from rhoai)
Parameters:
name: the InferenceService name - REQUIREDnamespace: user-specified namespace - REQUIREDverbosity: "full" - REQUIREDIf rhoai unavailable or returns error: Use resources_get (from openshift) with apiVersion: serving.kserve.io/v1beta1, kind: InferenceService, name: [name], namespace: [namespace]. Extract status from .status.conditions.
Early endpoint check:
MCP Tool: get_model_endpoint (from rhoai)
name: the InferenceService name, namespace: user-specified namespaceIf rhoai unavailable or returns error: Extract endpoint from .status.url of the InferenceService obtained via resources_get (from openshift).
An empty or error URL indicates deployment issues. Report endpoint availability status.
Present status conditions:
| Condition | Status | Reason | Message |
|---|---|---|---|
| Ready | [True/False/Unknown] | [reason] | [message] |
| PredictorReady | [True/False/Unknown] | [reason] | [message] |
| IngressReady | [True/False/Unknown] | [reason] | [message] |
Quick Assessment: Based on conditions, provide initial assessment (e.g., "PredictorReady is False -- the model container is not running. Likely a pod-level issue.")
Ask: "Continue with deep analysis of events and pods? (yes/no)"
WAIT for user confirmation.
MCP Tool: events_list (from openshift)
Parameters:
namespace: user-specified namespace - REQUIREDFilter events related to the InferenceService name.
MCP Tool: pods_list (from openshift)
Parameters:
namespace: user-specified namespace - REQUIREDlabelSelector: "serving.kserve.io/inferenceservice=[isvc-name]" - REQUIREDPresent findings:
Events:
| Time | Type | Reason | Message |
|---|---|---|---|
| [time] | [Normal/Warning] | [reason] | [message] |
Predictor Pods:
| Pod | Status | Restarts | Node | GPU |
|---|---|---|---|---|
| [pod-name] | [status] | [count] | [node] | [gpu-count] |
Issues Found:
Ask: "Continue to view pod logs? (yes/no)"
WAIT for user confirmation.
MCP Tool: pods_log (from openshift)
Parameters:
namespace: user-specified namespace - REQUIREDname: predictor pod name from Step 3 - REQUIREDcontainer: "kserve-container" - REQUIRED (main serving container)If the container has restarted, also retrieve previous logs.
Present log analysis:
Log Analysis:
For NIM-specific deployments, also check:
If the error is unrecognized -> Trigger live doc lookup:
Ask: "Continue to check related resources and observability? (yes/no)"
WAIT for user confirmation.
Check ServingRuntime:
MCP Tool: resources_get (from openshift)
Parameters:
apiVersion: "serving.kserve.io/v1alpha1" - REQUIREDkind: "ServingRuntime" - REQUIREDnamespace: user-specified namespace - REQUIREDname: runtime name from the InferenceService spec - REQUIREDVerify the runtime exists and its model format matches the InferenceService.
For NIM deployments -- Check Account CR:
MCP Tool: resources_get (from openshift)
Parameters:
apiVersion: "nim.opendatahub.io/v1alpha1" - REQUIREDkind: "Account" - REQUIREDnamespace: user-specified namespace - REQUIREDname: "nim-account" - REQUIREDIf ai-observability MCP is available:
get_deployment_info: Check if the model appears in monitoring and its initialization statusanalyze_vllm: Analyze performance metrics for slow inference (latency, throughput, errors, token rates)chat_vllm: Ask follow-up questions about analyzed metrics (e.g., "Why is latency spiking?")analyze_openshift with category "GPU & Accelerators": Check GPU health and utilizationquery_tempo_tool: Trace request latency if the symptom is slow responsesget_trace_details_tool: Drill into a specific trace ID to see span-level timingexecute_promql: Run custom PromQL queries for deeper metric investigation (e.g., vllm:request_success:ratio, GPU memory utilization)korrel8r_get_correlated: Correlate signals across the inference stack -- find related logs, traces, metrics, and alerts for the failing pod/namespace (query example: k8s:Pod:{"namespace":"[ns]","name":"[pod-name]"}, goals: ["log:application", "metric:metric", "trace:span"])If ai-observability not available: Skip with note: "Observability analysis skipped (ai-observability MCP not configured)."
Present findings:
Ask: "Continue to diagnosis summary? (yes/no)"
WAIT for user confirmation.
Present a structured diagnosis:
## Diagnosis Summary: [isvc-name]
### Root Cause
**Primary Issue:** [Categorized root cause]
| Category | Status | Details |
|----------|--------|---------|
| ServingRuntime | [OK/FAIL] | [details] |
| Pod Scheduling | [OK/FAIL] | [details] |
| Container Start | [OK/FAIL] | [details] |
| Model Loading | [OK/FAIL] | [details] |
| GPU Access | [OK/FAIL] | [details] |
| Endpoint Health | [OK/FAIL] | [details] |
### Evidence
- [Evidence 1 from events/logs/status]
- [Evidence 2]
### Recommended Actions
1. **[Action 1]** - [description]
2. **[Action 2]** - [description]
3. **[Action 3]** - [description]
### Verification Steps
After applying fixes:
1. Check InferenceService status: `resources_get` for the InferenceService
2. Verify pod is running: `pods_list` with label selector
3. Test endpoint: curl command to the inference URLEnd with options:
Would you like me to:
1. Execute a recommended fix
2. Dig deeper into a specific area
3. Debug a related resource (ServingRuntime, pod, NIM Account)
4. Invoke /serving-runtime-config to fix runtime issues
5. Exit debuggingWAIT for user to select next action.
For common issues (GPU scheduling, OOMKilled, image pull errors, RBAC), see common-issues.md.
Error: Pod logs show "Access Denied" or "NoSuchBucket" when loading model weights
Cause: S3 credentials are missing, expired, or the bucket/path is incorrect.
Solution:
storageUri in the InferenceService specError: NIM pod logs show NGC authentication failure, or TensorRT engine fails to compile for the available GPU
Cause: NGC API key is invalid/expired, or the GPU type is not supported by the NIM model profile.
Solution:
resources_get for accounts.nim.opendatahub.io/nim-setup to refresh credentials if expiredSee Prerequisites for the complete list of required and optional MCP tools.
/model-deploy - Redeploy or modify the InferenceService after fixing issues/serving-runtime-config - Fix or create ServingRuntime if runtime is the issue/nim-setup - Re-run NIM platform setup if NIM credentials are the issue/model-monitor - Check if TrustyAI monitoring detected issues before they became failuresSee skill-conventions.md for general HITL and security conventions.
Skill-specific checkpoints:
e46c4fa
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.