Configure TrustyAI Guardrails Orchestrator for LLM input/output content safety on OpenShift AI. Use when: - "Add guardrails to my LLM endpoint" - "Set up content safety for my model" - "Configure PII detection on my inference endpoint" - "Block prompt injection attacks" - "I need a guarded endpoint for my deployed model" Handles GuardrailsOrchestrator CR deployment, detector configuration (content safety, PII, prompt injection, toxicity), orchestration policies, and guarded endpoint validation. NOT for deploying models (use /model-deploy first). NOT for bias/drift monitoring (use /model-monitor). NOT for infrastructure observability (use /ai-observability).
69
85%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
View guide
Passed
No findings from the security scan
Required MCP Server: openshift (OpenShift MCP Server)
Required MCP Tools (from openshift):
resources_get (from openshift) - Get GuardrailsOrchestrator CR status, ConfigMapsresources_list (from openshift) - Check GuardrailsOrchestrator CRD availabilityresources_create_or_update (from openshift) - Create/update GuardrailsOrchestrator CR, detector ConfigMapsresources_delete (from openshift) - Remove detector configurations (with user confirmation)pods_list (from openshift) - Verify orchestrator and detector pods are runningpods_log (from openshift) - Retrieve orchestrator pod logs for troubleshootingevents_list (from openshift) - Check events for deployment issuesPreferred MCP Server: rhoai (RHOAI MCP Server) — used when available, automatic OpenShift fallback on failure
Preferred MCP Tools (from rhoai):
list_inference_services - List deployed models to identify guardrail targetsget_inference_service - Get InferenceService details (endpoint, runtime, status)get_model_endpoint - Get the model endpoint URL for orchestrator routingtest_model_endpoint - Test guarded endpoint after configurationdeploy_model - Deploy detector models (HuggingFace classifiers used as detectors)list_serving_runtimes - List runtimes for detector model deploymentrecommend_serving_runtime - Recommend runtime for detector modelsOptional MCP Server: ai-observability (AI Observability MCP)
Optional MCP Tools (from ai-observability):
execute_promql - Query guardrails metrics (request counts, block rates)analyze_vllm - Verify guarded endpoint performance impactCommon prerequisites (KUBECONFIG, OpenShift+RHOAI cluster, KServe, verification protocol): See skill-conventions.md.
Fallback templates: See openshift-fallback-templates.md for OpenShift YAML templates used when RHOAI tools are unavailable.
Additional cluster requirements:
/model-deploy)Use this skill when you need to:
Do NOT use this skill when:
/model-deploy)/model-monitor)/ai-observability)/debug-inference)MCP Tool: resources_list (from openshift)
Parameters:
apiVersion: "apiextensions.k8s.io/v1" - REQUIREDkind: "CustomResourceDefinition" - REQUIREDCheck for guardrailsorchestrators.trustyai.opendatahub.io CRD. This is a hard prerequisite — nothing in this skill works without it.
Error Handling:
spec.components.trustyai.managementState: Managed). Offer options: (1) Show enablement instructions, (2) Abort. WAIT for user decision.Ask the user for:
If user is unsure about target model, use list_inference_services (from rhoai) to present available models.
MCP Tool: list_inference_services (from rhoai)
Parameters:
namespace: user-specified namespace - REQUIREDverbosity: "standard" - OPTIONALIf rhoai unavailable or returns error: Use resources_list (from openshift) with apiVersion: serving.kserve.io/v1beta1, kind: InferenceService, namespace: [namespace].
Verify the selected InferenceService is Ready:
MCP Tool: get_inference_service (from rhoai)
Parameters:
name: selected InferenceService name - REQUIREDnamespace: target namespace - REQUIREDverbosity: "full" - REQUIREDIf rhoai unavailable or returns error: Use resources_get (from openshift) with apiVersion: serving.kserve.io/v1beta1, kind: InferenceService, name: [name], namespace: [namespace]. Extract status from .status.conditions.
If not Ready: Warn user and offer options: (1) Proceed anyway, (2) Invoke /debug-inference, (3) Abort. WAIT for user decision.
MCP Tool: get_model_endpoint (from rhoai)
Parameters:
name: selected InferenceService name - REQUIREDnamespace: target namespace - REQUIREDIf rhoai unavailable or returns error: Extract endpoint from .status.url of the InferenceService obtained via resources_get (from openshift).
Store the endpoint URL for orchestrator routing. Present configuration summary for confirmation. WAIT for user to confirm or modify.
Document Consultation (read before configuring detectors):
For each selected detector type:
Recommended model: ibm-granite/granite-guardian-3.1-2b (1 GPU, ~8Gi memory) per guardrails-detectors-reference.md.
Check if a compatible detector model is already deployed using list_inference_services (from rhoai). If one exists, offer to reuse it. WAIT for user decision.
If rhoai unavailable or returns error: Use resources_list (from openshift) with apiVersion: serving.kserve.io/v1beta1, kind: InferenceService, namespace: [namespace] to check for existing detector models. To list available runtimes when list_serving_runtimes is unavailable: Use resources_list (from openshift) with apiVersion: serving.kserve.io/v1alpha1, kind: ServingRuntime, namespace: [namespace].
If deploying a new detector:
MCP Tool: deploy_model (from rhoai)
Parameters:
name: "[isvc-name]-content-detector" (derived from target model name to avoid collisions) - REQUIREDnamespace: target namespace - REQUIREDruntime: appropriate runtime from list_serving_runtimes - REQUIREDmodel_format: "vLLM" - REQUIREDstorage_uri: "hf://ibm-granite/granite-guardian-3.1-2b" - REQUIREDgpu_count: 1 - OPTIONALmemory_request: "8Gi" - OPTIONALIf rhoai unavailable or returns error: Use resources_create_or_update (from openshift) to create the detector InferenceService CR directly with apiVersion: serving.kserve.io/v1beta1, kind: InferenceService.
Ask: "Deploy the content safety detector model? This creates an additional InferenceService. (yes/no/use-existing)"
WAIT for explicit confirmation. Monitor deployment until Ready.
Error Handling:
/debug-inference for the detector InferenceServiceUses built-in regex-based detectors (no model deployment needed). Generate appropriate regex patterns. Present patterns to user for review. WAIT for user decision.
For model-based detection: reuse the granite-guardian model from Step 3a (covers prompt injection). For keyword-based detection: configure patterns. WAIT for user decision.
Collect pattern name, regex, scope, and action from user.
Construct ConfigMap using the orchestrator config structure from guardrails-detectors-reference.md. Populate with detector configs from Step 3 and target model endpoint from Step 1.
MCP Tool: resources_create_or_update (from openshift)
Parameters:
manifest: ConfigMap YAML manifest as JSON string - REQUIREDConfigMap name: guardrails-config-[isvc-name]. Labels: app.kubernetes.io/part-of: trustyai-guardrails, trustyai.opendatahub.io/target-model: [isvc-name].
Display the full ConfigMap to user. Ask: "Proceed with this guardrails configuration? (yes/no/modify)"
WAIT for explicit confirmation.
Construct GuardrailsOrchestrator manifest using CRD spec from guardrails-detectors-reference.md. Key values: name=guardrails-[isvc-name], orchestratorConfig=guardrails-config-[isvc-name], enableBuiltInDetectors=true, enableGuardrailsGateway=true.
MCP Tool: resources_create_or_update (from openshift)
Parameters:
manifest: GuardrailsOrchestrator YAML manifest as JSON string - REQUIREDDisplay manifest to user. Ask: "Deploy this GuardrailsOrchestrator? (yes/no/modify)"
WAIT for explicit confirmation.
Error Handling:
MCP Tool: pods_list (from openshift)
Parameters:
namespace: target namespace - REQUIREDlabelSelector: "app.kubernetes.io/name=guardrails-[isvc-name]" - REQUIREDVerify orchestrator pod is Running. Poll every 15 seconds for up to 5 minutes.
On failure: Use pods_log and events_list (from openshift) to diagnose. Present options: (1) View full logs, (2) Check events, (3) Delete and recreate, (4) Abort. WAIT for user decision. NEVER auto-delete GuardrailsOrchestrator.
Get guarded endpoint: Use resources_get (from openshift) to read the GuardrailsOrchestrator CR status (apiVersion: trustyai.opendatahub.io/v1alpha1, kind: GuardrailsOrchestrator) and extract the guarded endpoint URL.
First, verify the original model still responds correctly:
MCP Tool: test_model_endpoint (from rhoai)
Parameters:
name: the original InferenceService name - REQUIREDnamespace: target namespace - REQUIREDIf rhoai unavailable or returns error: Note that test_model_endpoint only checks reachability, not actual inference. For a real inference test, use an in-cluster curl command: curl -X POST [endpoint]/v1/completions -H 'Content-Type: application/json' -d '{"model":"[model]","prompt":"Hello","max_tokens":10}'
Then test the guarded endpoint directly. The guarded endpoint is a different URL from the original — obtain it from the GuardrailsOrchestrator CR status (Step 6). If the guarded endpoint is only available cluster-internally, set up port-forwarding to the orchestrator service first:
oc port-forward svc/guardrails-[isvc-name] 8080:8080 -n [namespace]Run a safe request against the guarded endpoint to confirm it proxies correctly, then run an unsafe request (e.g., prompt injection attempt) to verify the detectors are active. Present both results to the user with pass/fail for each test.
Present summary showing: guarded vs original endpoint URLs, active detectors table (name, type, scope, policy), usage instructions (applications should use guarded endpoint), and next steps (/model-monitor, /ai-observability).
For common issues (GPU scheduling, OOMKilled, image pull errors, RBAC), see common-issues.md.
Error: Content safety or prompt injection detector model InferenceService fails to start
Cause: Insufficient resources (GPU/memory) for the detector model, or runtime compatibility issues.
Solution:
/debug-inference to troubleshoot the detector InferenceServiceError: Requests to the guarded endpoint return 502 Bad Gateway or 503 Service Unavailable
Cause: The orchestrator cannot reach the underlying model endpoint, or the detector service is down.
Solution:
test_model_endpoint from rhoaipods_logorchestrator.target_model.endpoint URL is correctError: Cannot create guardrailsorchestrators resource — 403 Forbidden
Cause: The user lacks RBAC for the GuardrailsOrchestrator CRD, which is typically cluster-admin only.
Solution: Provide the user with the complete GuardrailsOrchestrator CR YAML and instruct them to ask a cluster administrator to apply it. The detectors ConfigMap (which only requires namespace edit role) can still be created by the skill.
Error: Guarded endpoint is significantly slower than direct endpoint, or legitimate requests are blocked
Cause: Too many model-based detectors add latency; overly aggressive thresholds or broad regex patterns cause false positives.
Solution:
/guardrails-config with modified configurationSee Prerequisites for the complete list of required and optional MCP tools.
/model-deploy - Deploy the target LLM before configuring guardrails; also used to deploy detector models/model-monitor - Add bias and drift monitoring (complements safety guardrails)/debug-inference - Troubleshoot failed detector model deployments or guarded endpoint issues/ai-observability - Monitor guardrails impact on latency and throughput/serving-runtime-config - Configure custom runtime for detector models if neededSee skill-conventions.md for general HITL and security conventions.
Skill-specific checkpoints:
e46c4fa
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.