The SRE team's runbook for triaging production latency and error-rate incidents. Use this whenever investigating an incident, a latency spike, elevated error rates, or when asked "what caused X" about a production service.
69
84%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
View guide
Passed
No findings from the security scan
If you change the order below, say why in #sre.
p99_latency_ms / error_rate for the paged service. State the gap ("deploy 14:31, p99 moves 14:33").db_pool_utilization across checkout/cart/auth/inventory, then upstream deps.In rough order of how often:
One line at the bottom:
Root cause:
<sha>— one sentence on the mechanism.
If it wasn't a deploy, put the component or upstream dep where the sha goes (db-primary, stripe-api, whatever). Still one sentence.
Everything above that line is evidence. Keep it short; the long version goes in the postmortem doc.
068b84b
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.