Skip to content

Operational Excellence

Run systems professionally throughout their full lifecycle.


Principles

  • Ownership — deployment is not the finish line
  • Automation By Default — manual processes do not scale
  • Continuous Improvement — every incident is a learning opportunity
  • Document critical knowledge — knowledge trapped in people creates organisational risk

Service Ownership

Every service should have: - A named owner - Up-to-date documentation - Monitoring in place - A runbook for common failure scenarios


Monitoring and Observability

→ Monitoring and Observability

The four golden signals: - Latency — how long requests take - Traffic — volume of demand - Errors — rate of failures - Saturation — resource consumption

Observability stack: - Logs — structured and searchable - Metrics — quantitative measurements over time - Traces — request journeys across services


Incident Management

→ Incident Management

Severity Description
Sev 1 Critical outage
Sev 2 Major degradation
Sev 3 Limited impact
Sev 4 Minor issue

Post-incident reviews focus on root causes and system improvements, not blame.


Strong Signals

  • Automation covers repetitive operational tasks
  • Runbooks exist and are kept current
  • Incidents are detected quickly and resolved consistently
  • Self-service capabilities reduce bottlenecks

Weak Signals

  • Tribal knowledge and key-person dependency
  • Manual interventions for routine tasks
  • Operational bottlenecks
  • Systems with no named owner

← Engineering Excellence