Manages production incident response workflows from initial triage through postmortem. Guides responders through alert acknowledgment, mitigation, status communication, and follow-up. Use when the user mentions outages, service degradation, on-call pages, incident response, postmortems, error rate spikes, latency issues, or on-call escalation. Covers rollback decisions, stakeholder updates, and structured post-incident review.
72
87%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
View guide
Passed
No findings from the security scan
Use this skill when a service is down or behaving badly and the team needs to coordinate a response.
When an alert fires or a customer reports an issue:
#inc-<short-name>) and post a one-line summaryOnce you have at least one responder online, focus on stopping the bleeding:
Don't worry about root cause yet. Mitigation comes first.
Post status updates every 15 minutes in the incident channel and on the public status page. Use this template (save to a shared runbook file such as STATUS_UPDATE_TEMPLATE.md for team customization):
When the incident is contained, do stuff to follow up. Things like writing a postmortem, scheduling a review, etc.
Immediate (within 1 hour of resolution):
Short-term (within 3 business days):
POSTMORTEM_TEMPLATE.md (create this file if it does not exist; see outline below)Postmortem outline (extract to POSTMORTEM_TEMPLATE.md for reuse):
Review meeting checklist:
73eda88
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.