Incident Management¶
Outages are inevitable. Extended outages are not.
The goal is fast detection, fast recovery, and learning — not blame.
Severity Levels¶
| Severity | Description | Example |
|---|---|---|
| Sev 1 | Critical outage — system unavailable | Production pipeline completely down |
| Sev 2 | Major degradation — significant impact | Key reports not refreshing |
| Sev 3 | Limited impact — partial or minor issue | One dashboard slow, others fine |
| Sev 4 | Minor issue — minimal impact | Non-critical alert firing spuriously |
Response Process¶
- Detect — alert fires or issue is reported
- Acknowledge — someone takes ownership
- Assess — determine severity and impact
- Communicate — notify stakeholders appropriate to severity
- Mitigate — stop the bleeding, restore service
- Resolve — fix the underlying cause
- Review — post-incident review
Post-Incident Review¶
Run a PIR for every Sev 1 and Sev 2. Optional for Sev 3.
Focus on: - What happened and when (timeline) - Root cause - What helped, what hindered the response - System improvements to prevent recurrence - Process improvements
Avoid blame. Focus on systems, not individuals.
On-Call¶
- Every critical system should have a defined on-call owner
- Escalation paths should be documented before they are needed
- On-call burden should be shared fairly across the team