Fix a failed Orchestra pipeline once the failure has been identified as an Orchestra-platform / configuration issue — pipeline YAML misconfiguration, wrong or missing task inputs, task ordering, env/connection wiring, a transient Orchestra-side blip needing a plain retry, or an Orchestra-backed pipeline that needs update_pipeline. Also the fallback fixer for repo-level code fixes in integrations that don't have a dedicated skill (e.g. a Snowflake/HTTP SQL bug needing a PR). Normally invoked by identify-pipeline-error after it classifies the cause; for dbt-code or Python-code failures, that router calls fix-pipeline-dbt-task or fix-pipeline-python-task instead. This skill is the FIX half (apply fix → PR/poll → retry → confirm → optionally remember); identification and classification live in identify-pipeline-error.
79
90%
Does it follow best practices?
Impact
100%
1.00xAverage score across 1 eval scenario
Low
Low-risk findings worth noting
The canonical home for this skill is orchestra-hq/orchestra-skills
Fix and retry an already-diagnosed Orchestra pipeline — and optionally remember what worked.
This skill assumes the Orchestra MCP server is connected. All MCP calls are scoped to the user's workspace.
Use Orchestra MCP tools for all operations in this skill (read a pipeline's full definition with
get_pipeline). Argument summaries: ../../references/orchestra/mcp/tools-quick-ref.md.
identify-pipeline-error is the entry point that parses the user's input (URL, UUID, alias, error
text, Slack alert), finds the failed run and task, and classifies the cause. It hands you:
pipelineId, pipelineName, envName, run branch/commit, triggeredByid, taskName, taskId, integration, integrationJob, message,
externalStatus, externalMessage, taskParameters, runParameters, connectionId,
numberOfAttemptsUse these directly and proceed to Step 3. If invoked standalone (no handoff), run
identify-pipeline-error first to establish the failed run/task and the cause, then come back here.
Execute these steps in order. Each step feeds the next.
Goal: Get the raw evidence — logs, artifacts, and operations.
For each failed task run:
Logs: Call list_task_run_logs to list available log files. Then fetch each log
with download_task_run_log. Focus on the last ~256KB of large logs using
range_header (for example bytes=-262144).
Artifacts: Call list_task_run_artifacts. For dbt tasks, look for
run_results.json and manifest.json. Download relevant artifacts with
download_task_run_artifact.
Operations: Call list_operations filtered by task_run_id to see
sub-operations (individual dbt models, Snowflake queries, etc.) and their statuses.
Read ../../references/orchestra/pipeline/diagnosis-patterns.md before proceeding to Step 4. It contains
integration-specific error patterns that will help classify the failure.
Goal: Pin down the specific root cause so the fix is exact. identify-pipeline-error already
classified this as an Orchestra-platform/config issue (or a repo-level code fix with no dedicated
skill) and routed you here — you are not re-deciding the category. Using the evidence from Step 3
plus ../../references/orchestra/pipeline/diagnosis-patterns.md:
Identify the root cause precisely. Not just "config error" but "task load_events is missing
the dbt_branch run input, so it cloned the default branch" or "column user_email does not
exist in analytics.users — a schema migration removed it." Confirm it's actionable as a
YAML/config change, a retry, or a repo PR — i.e. genuinely an Orchestra-platform/config or
repo-code fix.
Scope check — hand back if it isn't yours. If the evidence now points to a dbt code
failure, a Python code failure, or a category identify-pipeline-error handles itself (data
quality, vendor/ingestion needing a UI fix, auth, network, pure upstream), stop and hand back
to identify-pipeline-error so it can route correctly — don't force a fix here.
(Optional) Recall past fixes. Past-fix memory is deferred to the calling agentic client.
If your client exposes persistent memory (e.g. Claude Code memory, Cursor rules/memories), check
it for similar past fixes. As a fallback, read a local
../../references/orchestra/pipeline/knowledge-store.md if the user keeps one — it ships empty
and may not exist. Optional: skip when no memory is available. Treat any recalled entry as
historical context and re-verify it still applies before acting on it.
Present the root cause clearly: specific cause, evidence (which log line / error / operation failed), and confidence (high/medium/low).
Read ../../references/orchestra/pipeline/remediation-playbooks.md before proceeding to Step 5.
Goal: Fix the issue or tell the user exactly what to do.
Based on the diagnosis, consult ../../references/orchestra/pipeline/remediation-playbooks.md and take action:
Fixes the agent can apply directly:
start_pipelineupdate_pipeline to fix configuration errors like wrong parameters, missing
environment variables, or incorrect task orderingrun_inputs in start_pipeline to override
problematic input valuesgh CLI to create a pull request with the fix directly.
Do not ask the user to make the change themselves. Workflow:
fix/missing-pyproject-toml)gh pr create targeting the failing branchFixes that require user action (explain clearly — but still poll after PR if one was opened):
Always explain what you're doing and why before taking action.
Goal: Watch the PR and trigger the pipeline rerun once it merges — without making the user babysit it.
After sharing the PR URL, emit one status line and keep watching the PR until it reaches a terminal state (or the user asks you to stop):
⏳ PR #178 open — checking every 60 s; will trigger the pipeline on merge.Polling loop:
Check PR state:
gh pr view {pr_number} --repo {owner/repo} --json state,mergedAtIf state == "MERGED": Proceed immediately to Step 6 — trigger start_pipeline
using the original pipeline ID and environment. No confirmation needed (the user
already approved the fix by merging the PR).
If state == "CLOSED" (not merged): The PR was closed without merging. Report
this and ask the user how to proceed — do not auto-retry.
If state == "OPEN": Wait ~60 seconds, then check again; after several checks with
no merge, widen the interval to a few minutes. Use whatever scheduling mechanism your
client provides — if it can re-invoke you on a timer, schedule the next check and hand
back control; otherwise keep polling in the same conversation. Either way, retain the PR
number, repo, pipeline ID, and environment so each check resumes the same fix workflow.
Polling output format (one line per check, not a full summary):
⏳ PR #178 — OPEN (2 min elapsed, next check in 60 s)
⏳ PR #178 — OPEN (3 min elapsed, next check in 60 s)
✅ PR #178 — MERGED — triggering pipeline rerun…Do not re-diagnose or re-explain the fix on each poll tick. One line only.
Goal: Confirm the fix worked.
This step is entered either (a) directly after a non-PR fix, or (b) automatically from Step 5b once the PR is merged.
start_pipeline
run_inputs if applicableget_pipeline_run_status every ~30 secondsPersisting fixes is optional and deferred to the calling agentic client. Only do it when the user wants a durable record — it is off by default, and nothing workspace-specific should be committed to this repository.
../../references/orchestra/pipeline/knowledge-store.md,
append an entry using the template at the bottom of that file. The published file ships empty.When you do record a fix, capture: date, pipeline name, error category, integration, root cause, fix applied, and whether the first diagnosis was correct.
If you discover a genuinely new, generic diagnosis pattern, consider noting it in
../../references/orchestra/pipeline/diagnosis-patterns.md — but keep workspace-specific detail
(pipeline IDs, connection names, account identifiers) out of shared reference files.
Be succinct. Users are engineers dealing with broken pipelines — they want facts and actions, not explanations. No preamble, no summaries of what you just did, no "great news". If the answer fits in one line, use one line. Cut any sentence that doesn't add new information.
All user-facing output must follow these templates exactly. Consistent structure makes it easy
to scan at a glance. (Multi-pipeline "what's broken" triage is owned by identify-pipeline-error,
not this skill.)
Use a consistent header and structured block every time, then evidence and confidence.
## Root cause: `<pipeline name>`
**Root cause:** One specific sentence — not "query error" but exactly which object/column/table/input.
**Integration:** DBT_CORE / SNOWFLAKE / HTTP (or similar)
**Connection:** connection-id-here
**Confidence:** High / Medium / Low
**Evidence:**
- Exact log line or error message (quoted)
- Which operation or model failed
- Any corroborating signals (exit code, attempt count, etc.)Always present options as a numbered list with a clear owner label on each:
**Fix options:**
1. **[Agent can apply]** Short description of what will be done and why it fixes the issue.
2. **[Agent opens a PR]** For code changes in Git-backed pipelines — describe what file/change
will be committed and which branch the PR targets.
3. **[User action needed]** Only for things the agent genuinely cannot do: credential rotation,
firewall changes, permission grants. Include the specific UI path, command, or SQL.
4. **[Needs more info]** What you'd need to know to proceed (e.g. "Does the EVENTS table
exist? Run: SELECT COUNT(*) FROM SNOWFLAKE_WORKING.PUBLIC.EVENTS").
Recommended: Option N — one sentence on why this is the right call.Never present options without a recommendation. Never use vague labels like "you could try".
Never label a code fix as [User action needed] — open the PR instead.
During polling, emit a single line per status check. Do not repeat the full context.
▶ Run e2049b86 — RUNNING (0:42 elapsed)
▶ Run e2049b86 — RUNNING (1:30 elapsed)
✅ Run e2049b86 — SUCCEEDED (2:15 elapsed)On failure during retry, immediately switch to diagnosis format (don't just say "it failed").
Always end a successful fix with a compact summary block:
## Fixed: `<pipeline name>`
- **Run:** `<new-run-id>` — SUCCEEDED
- **Root cause:** (one line)
- **Fix applied:** (one line)
- **Duration:** X min
- **Recorded:** Saved to client memory ✓ (only if the user opted in — omit this line otherwise)update_pipeline.
Git-backed pipelines require code changes committed to the repository. Check storageProvider
in the list_pipelines response to determine which type it is.3a29fe4
Canonical home
since Jun 16, 2026
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.