Manages production incident response workflows from initial triage through postmortem. Guides responders through alert acknowledgment, mitigation, status communication, and follow-up. Use when the user mentions outages, service degradation, on-call pages, incident response, postmortems, error rate spikes, latency issues, or on-call escalation. Covers rollback decisions, stakeholder updates, and structured post-incident review.
71
89%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
Passed
No findings from the security scan
Use this skill when a service is down or behaving badly and the team needs to coordinate a response.
When an alert fires or a customer reports an issue:
#inc-<short-name>) and post a one-line summaryOnce you have at least one responder online, focus on stopping the bleeding:
Don't worry about root cause yet. Mitigation comes first.
Post status updates every 15 minutes in the incident channel and on the public status page. Use this template (save to a shared runbook file such as STATUS_UPDATE_TEMPLATE.md for team customization):
When the incident is contained, do stuff to follow up. Things like writing a postmortem, scheduling a review, etc.
Immediate (within 1 hour of resolution):
Short-term (within 3 business days):
POSTMORTEM_TEMPLATE.md (create this file if it does not exist; see outline below)Postmortem outline (extract to POSTMORTEM_TEMPLATE.md for reuse):
Review meeting checklist:
73eda88
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.