Operational Excellence¶
Run systems professionally throughout their full lifecycle.
Principles¶
- Ownership — deployment is not the finish line
- Automation By Default — manual processes do not scale
- Continuous Improvement — every incident is a learning opportunity
- Document critical knowledge — knowledge trapped in people creates organisational risk
Service Ownership¶
Every service should have: - A named owner - Up-to-date documentation - Monitoring in place - A runbook for common failure scenarios
Monitoring and Observability¶
→ Monitoring and Observability
The four golden signals: - Latency — how long requests take - Traffic — volume of demand - Errors — rate of failures - Saturation — resource consumption
Observability stack: - Logs — structured and searchable - Metrics — quantitative measurements over time - Traces — request journeys across services
Incident Management¶
| Severity | Description |
|---|---|
| Sev 1 | Critical outage |
| Sev 2 | Major degradation |
| Sev 3 | Limited impact |
| Sev 4 | Minor issue |
Post-incident reviews focus on root causes and system improvements, not blame.
Strong Signals¶
- Automation covers repetitive operational tasks
- Runbooks exist and are kept current
- Incidents are detected quickly and resolved consistently
- Self-service capabilities reduce bottlenecks
Weak Signals¶
- Tribal knowledge and key-person dependency
- Manual interventions for routine tasks
- Operational bottlenecks
- Systems with no named owner