Create, run, schedule, and monitor Data Science Pipelines (Kubeflow Pipelines 2.0) on OpenShift AI. Use when: - "Run a pipeline in my project" - "Schedule a recurring pipeline" - "Check my pipeline run status" - "List pipeline runs and their logs" - "Set up the pipeline server" - "Delete a pipeline or pipeline run" Handles pipeline server setup, pipeline run submission from YAML, scheduling recurring runs, monitoring execution, and viewing step logs. NOT for creating data science projects (use /ds-project-setup). NOT for deploying models (use /model-deploy). NOT for model training jobs (use training skills).
72
90%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
View guide
Low
Low-risk findings worth noting
Create, run, schedule, and monitor Data Science Pipelines (Kubeflow Pipelines 2.0) on Red Hat OpenShift AI. Handles the full pipeline lifecycle: verifying or setting up the pipeline server (DSPA), submitting pipeline runs from YAML definitions, scheduling recurring runs with cron expressions, monitoring run status with step-level progress, viewing pipeline step logs, and deleting pipeline resources with proper warnings.
Required MCP Server: openshift (OpenShift MCP Server)
Required MCP Tools (from openshift):
resources_create_or_update - Create DSPA CR (pipeline server), PipelineRun and ScheduledWorkflow CRsresources_list - List PipelineRun resources, DSPA statusresources_get - Get PipelineRun status, DSPA detailsresources_delete - Delete pipeline run resources, DSPAevents_list - Check pipeline pod events for errorspods_list - List pipeline step podspods_log - Retrieve pipeline step container logsPreferred MCP Server: rhoai (RHOAI MCP Server) — used when available, automatic OpenShift fallback on failure
Preferred MCP Tools (from rhoai):
list_data_science_projects - Validate namespace is an RHOAI Data Science Projectget_pipeline_server - Check pipeline server (DSPA) status in a projectdelete_pipeline_server - Delete pipeline server and all pipeline infrastructurelist_resources - List pipeline resources in a namespace (resource_type="pipelines")get_resource - Get pipeline resource details (resource_type="pipeline")list_resource_names - Token-efficient pipeline name listing (resource_type="pipelines")resource_status - Quick pipeline status check (resource_type="pipeline")diagnose_resource - Full diagnostic for a pipeline (resource_type="pipeline")list_data_connections - Verify S3 data connections for pipeline artifact storageproject_summary - Project overview including pipeline statusCommon prerequisites (KUBECONFIG, OpenShift+RHOAI cluster, verification protocol): See skill-conventions.md.
Fallback templates: See openshift-fallback-templates.md for OpenShift YAML templates used when RHOAI tools are unavailable.
Additional cluster requirements:
opendatahub.io/dashboard: "true")Use this skill when you need to:
Do NOT use this skill when:
/ds-project-setup)/model-deploy)/model-registry)Ask the user what they want to do: Setup server, List pipelines/runs, Run pipeline, Schedule recurring run, Monitor run, View logs, Delete resources.
Ask for target namespace. Validate via list_data_science_projects (from rhoai). If invalid, suggest /ds-project-setup.
If rhoai unavailable or returns error: Use resources_list (from openshift) with apiVersion: v1, kind: Namespace, labelSelector: opendatahub.io/dashboard=true.
Route: Setup -> Step 2, List -> Step 3, Run -> Step 4, Schedule -> Step 5, Monitor -> Step 6, Logs -> Step 7, Delete -> Step 8.
Check via get_pipeline_server (from rhoai) with namespace. If healthy, proceed. If unhealthy, offer diagnostics via diagnose_resource. If not exists, offer setup.
If rhoai unavailable or returns error: Use resources_get (from openshift) with apiVersion: datasciencepipelinesapplications.opendatahub.io/v1alpha1, kind: DataSciencePipelinesApplication, name: dspa, namespace: [namespace]. Check .status.conditions for Ready=True.
For setup: Check data connections via list_data_connections (from rhoai). If none exist, offer to delegate to /ds-project-setup.
If rhoai unavailable or returns error: Use resources_list (from openshift) with apiVersion: v1, kind: Secret, namespace: [namespace], labelSelector: opendatahub.io/dashboard=true. Filter by annotation opendatahub.io/connection-type: s3.
Gather: Select from available data connections (the data connection name is the S3 secret name). The bucket, endpoint, and region can be extracted from the data connection secret. Present configuration for review. WAIT for confirmation.
Pipeline Server Creation (OpenShift direct — create_pipeline_server from rhoai is not used because it constructs invalid DSPA manifests):
MCP Tool: resources_create_or_update (from openshift)
Create a DataSciencePipelinesApplication CR. See openshift-fallback-templates.md for the YAML template.
Parameters to fill in the template:
namespace: target namespacebucket: S3 bucket name from the data connectionhost: S3 endpoint without protocol prefix (e.g., minio.namespace.svc:9000)scheme: http or httpssecretName: name of the S3 data connection secretregion: AWS region or empty string for MinIOVerify DSPA is ready:
MCP Tool: resources_get (from openshift)
apiVersion: datasciencepipelinesapplications.opendatahub.io/v1alpha1, kind: DataSciencePipelinesApplication, name: dspa, namespace: [namespace]Check .status.conditions for Ready=True. Poll every 15 seconds until ready or timeout (5 minutes).
Error Handling:
/ds-project-setupUse list_resources (from rhoai) with resource_type="pipelines", namespace for detailed listing, or list_resource_names for a quick name-only view.
For specific run status: resource_status (from rhoai) with resource_type="pipeline", name, namespace.
For project-wide overview: project_summary (from rhoai) with namespace.
If rhoai unavailable or returns error: Use resources_list (from openshift) with apiVersion: tekton.dev/v1, kind: PipelineRun, namespace: [namespace] to list pipeline runs directly.
If pipeline server not configured, suggest setup via Step 2.
Verify pipeline server is ready via get_pipeline_server (from rhoai). If not ready, offer Step 2.
Gather from user: pipeline definition (file path or inline YAML/JSON), run name (DNS-compatible), pipeline parameters (key-value pairs), service account (default: pipeline-runner-dspa).
Read pipeline definition using the Read tool if a file path is provided. Present configuration for review. WAIT for confirmation.
MCP Tool: resources_create_or_update (from openshift)
Parameters:
resource: PipelineRun CR (apiVersion: tekton.dev/v1, kind: PipelineRun) with spec.pipelineSpec or spec.pipelineRef, spec.params, spec.serviceAccountName - REQUIREDIf the pipeline uses KFP v2 compiled format (Argo-based), adapt apiVersion/kind accordingly.
Proceed to Step 6 to monitor the run.
Verify pipeline server is ready (same as Step 4).
Gather from user: pipeline reference (name or YAML), schedule (cron expression or natural language), pipeline parameters, max concurrent runs (default: 1), optional start/end time.
Convert natural language to cron if needed. Present schedule configuration for review. WAIT for confirmation.
MCP Tool: resources_create_or_update (from openshift)
Parameters:
resource: ScheduledWorkflow CR (apiVersion: scheduledworkflows.kubeflow.org/v1beta1, kind: ScheduledWorkflow) with spec.enabled, spec.maxConcurrency, spec.trigger.cronSchedule.cron, spec.trigger.cronSchedule.startTime/endTime, spec.workflow.spec.params - REQUIREDError Handling:
Get run status via get_resource (from rhoai) with resource_type="pipeline", name, namespace, verbosity="full".
For deeper diagnostics: diagnose_resource (from rhoai) with resource_type="pipeline", name, namespace.
If rhoai unavailable or returns error: Use resources_get (from openshift) with apiVersion: tekton.dev/v1, kind: PipelineRun, name: [run-name], namespace: [namespace]. Extract task status from .status.childReferences or .status.taskRuns.
Track step-level progress via resources_get (from openshift) with apiVersion tekton.dev/v1, kind PipelineRun. Extract task statuses from .status.childReferences or .status.taskRuns.
List pipeline pods via pods_list (from openshift) with namespace and labelSelector="tekton.dev/pipelineRun=<run-name>".
Present step progress table:
| Step | Status | Duration | Message |
|---|---|---|---|
| [step-name] | [Running/Succeeded/Failed] | [duration] | [message] |
On failure: Present options: (1) View step logs, (2) Check events, (3) Run diagnostics, (4) Retry. WAIT for user decision. NEVER auto-retry or auto-delete failed runs.
Error Handling:
From the PipelineRun status (Step 6), identify the pod name for the failing or target step. If user did not specify a step, list available steps and ask which one to inspect.
MCP Tool: pods_log (from openshift)
Parameters:
namespace: target namespace - REQUIREDname: pod name of the pipeline step - REQUIREDcontainer: container name - OPTIONAL (specify if multiple containers)Also check events_list (from openshift) filtered by namespace for scheduling/resource issues.
Suggest fixes based on common log patterns: OOMKilled -> increase memory limits; ImagePullBackOff -> verify image reference; AccessDenied on S3 -> check data connection via /ds-project-setup; Permission denied -> verify ServiceAccount permissions.
Error Handling:
Delete a pipeline run:
Get run details via get_resource (from rhoai) with resource_type="pipeline", name, namespace. Display run details (name, status, created, step count).
Ask: "Delete pipeline run <name>? This will remove the run record and associated pods. (yes/no)" WAIT for confirmation.
MCP Tool: resources_delete (from openshift) with apiVersion tekton.dev/v1, kind PipelineRun, name, namespace.
Delete the pipeline server (DESTRUCTIVE):
Display warning: deleting the DSPA removes all pipeline runs, history, API/UI endpoints, and terminates running pipelines. S3 data is preserved.
Ask: "Type the namespace name <namespace> to confirm deletion:" WAIT for typed confirmation. Verify exact match.
MCP Tool: delete_pipeline_server (from rhoai) with namespace, confirm=true.
Verify via get_pipeline_server (from rhoai). Confirm removal.
Error Handling:
For common issues (GPU scheduling, OOMKilled, image pull errors, RBAC), see common-issues.md.
Cause: Invalid S3 credentials, unreachable S3 endpoint, or database connectivity issues.
Solution: Use diagnose_resource for automated diagnostics. Check data connection credentials via list_data_connections. Verify S3 bucket accessibility. Check DSPA pod logs for specific errors.
Cause: Insufficient resources, missing ServiceAccount, unbound PVC, or scheduling issues.
Solution: Check events for scheduling errors. Verify ServiceAccount pipeline-runner-dspa exists. Check ResourceQuota/LimitRange. For GPU pipelines, verify availability via /ai-observability.
Cause: Step container exceeded memory limit.
Solution: View step logs before OOM kill. Increase resources.limits.memory in the pipeline YAML. Consider streaming data for data-intensive steps.
Cause: Expired credentials, missing bucket, or unreachable endpoint.
Solution: Verify data connection via list_data_connections. Check S3 bucket exists. Re-create data connection if credentials rotated.
Cause: Invalid cron, ScheduledWorkflow controller not running, or schedule disabled.
Solution: Verify ScheduledWorkflow CR spec.enabled is true. Validate cron format. Check controller pod logs. Verify startTime/endTime constraints.
See Prerequisites for the complete list of required MCP tools.
/ds-project-setup - Create a Data Science Project with data connections (prerequisite for pipeline server)/model-deploy - Deploy a model produced by a pipeline/model-registry - Register models produced by pipeline runs/ai-observability - Check GPU/cluster resources before running resource-intensive pipelines/debug-inference - Debug models deployed from pipeline outputsSee skill-conventions.md for general HITL and security conventions.
Skill-specific checkpoints:
User: "Set up the pipeline server in my ml-training project and run a data preprocessing pipeline"
Skill response: Validates project, sets up pipeline server with S3 data connection after confirmation, monitors readiness, then gathers pipeline YAML, presents run config for confirmation, submits PipelineRun, and monitors step-level progress.
e46c4fa
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.