Monitoring & Logging — Observability for Real Systems

Observability isn't a dashboard you build once. It's the set of signals that tell you something is wrong before users do, and the tools that let you understand why. This category covers practical observability for production Kubernetes systems — from choosing between ELK and Loki as your logging stack, to fixing liveness and readiness probe failures that cause silent traffic loss, to building runbooks that actually get used during incidents.

Prometheus Grafana Loki ELK Alertmanager Probes Metrics Logs SLOs Runbooks

Health Check & Probe Issues

Kubernetes probe failures that cause restarts or silent traffic loss.

Observability & Logging Guides

Architecture decisions and implementation guides for monitoring production systems.

Browse by topic

Every article on DevOps Compass is organized into a focused category.