Skip to content

Monitoring and Observability

If it matters, it should be monitored. You cannot improve what you cannot measure.


The Four Golden Signals

Signal What it measures
Latency How long requests or processes take
Traffic Volume of demand on the system
Errors Rate of failures and exceptions
Saturation Resource consumption — CPU, memory, disk, queue depth

Start here. These four signals surface most meaningful problems.


Observability Stack

Logs

  • Structured (JSON) and searchable
  • Include context: timestamp, service, correlation ID, severity
  • Avoid logging sensitive data

Metrics

  • Quantitative measurements over time
  • Use for alerting and dashboards
  • Examples: pipeline run duration, row counts, error rates

Traces

  • Request journeys across services
  • Useful for diagnosing latency in distributed systems

Alerting Principles

  • Alert on symptoms, not causes — alert when users are affected, not when an internal metric spikes
  • Every alert should have a runbook or clear response path
  • Reduce noise — alert fatigue leads to ignored alerts

Suggested Tooling

Tool Purpose
Datadog Metrics, logs, traces, dashboards
Grafana + Prometheus Open-source metrics and dashboards
CloudWatch AWS-native monitoring
OpenTelemetry Vendor-neutral instrumentation

← Operational Excellence ← Engineering Excellence