Monitoring and Observability¶
If it matters, it should be monitored. You cannot improve what you cannot measure.
The Four Golden Signals¶
| Signal | What it measures |
|---|---|
| Latency | How long requests or processes take |
| Traffic | Volume of demand on the system |
| Errors | Rate of failures and exceptions |
| Saturation | Resource consumption — CPU, memory, disk, queue depth |
Start here. These four signals surface most meaningful problems.
Observability Stack¶
Logs¶
- Structured (JSON) and searchable
- Include context: timestamp, service, correlation ID, severity
- Avoid logging sensitive data
Metrics¶
- Quantitative measurements over time
- Use for alerting and dashboards
- Examples: pipeline run duration, row counts, error rates
Traces¶
- Request journeys across services
- Useful for diagnosing latency in distributed systems
Alerting Principles¶
- Alert on symptoms, not causes — alert when users are affected, not when an internal metric spikes
- Every alert should have a runbook or clear response path
- Reduce noise — alert fatigue leads to ignored alerts
Suggested Tooling¶
| Tool | Purpose |
|---|---|
| Datadog | Metrics, logs, traces, dashboards |
| Grafana + Prometheus | Open-source metrics and dashboards |
| CloudWatch | AWS-native monitoring |
| OpenTelemetry | Vendor-neutral instrumentation |