Skip to content

Incident Management

Outages are inevitable. Extended outages are not.

The goal is fast detection, fast recovery, and learning — not blame.


Severity Levels

Severity Description Example
Sev 1 Critical outage — system unavailable Production pipeline completely down
Sev 2 Major degradation — significant impact Key reports not refreshing
Sev 3 Limited impact — partial or minor issue One dashboard slow, others fine
Sev 4 Minor issue — minimal impact Non-critical alert firing spuriously

Response Process

  1. Detect — alert fires or issue is reported
  2. Acknowledge — someone takes ownership
  3. Assess — determine severity and impact
  4. Communicate — notify stakeholders appropriate to severity
  5. Mitigate — stop the bleeding, restore service
  6. Resolve — fix the underlying cause
  7. Review — post-incident review

Post-Incident Review

Run a PIR for every Sev 1 and Sev 2. Optional for Sev 3.

Focus on: - What happened and when (timeline) - Root cause - What helped, what hindered the response - System improvements to prevent recurrence - Process improvements

Avoid blame. Focus on systems, not individuals.


On-Call

  • Every critical system should have a defined on-call owner
  • Escalation paths should be documented before they are needed
  • On-call burden should be shared fairly across the team

← Operational Excellence ← Engineering Excellence