Kubernetes errors happen fast and the documentation is rarely enough. This category covers real production failure scenarios — the exact commands to diagnose them, the logic behind each fix, and the patterns that prevent the same issue from coming back. Whether you're chasing a CrashLoopBackOff at 2am or tuning probe timing before a release, every guide here is built around how Kubernetes actually behaves, not how the spec says it should.
Troubleshooting
The most-searched Kubernetes errors, with step-by-step diagnosis and fixes.
Pods stuck at Init:0/1 or Init:CrashLoopBackOff — failing init containers, missing ConfigMaps/Secrets, volume waits, and dependency checks.
Diagnose and resolve image pull failures caused by registry credentials, wrong image tags, or network restrictions.
Step-by-step diagnosis for containers that keep crashing — from misconfigured env vars to missing secrets and bad entrypoints.
Diagnose Out of Memory kills — covers memory limits, JVM heap sizing, memory leaks, sidecar pressure, and VPA right-sizing.
Stop Kubernetes from killing healthy containers — covers initialDelaySeconds, timeoutSeconds, wrong paths, and startupProbe configuration.
Fix pods stuck at 0/1 Ready — covers dependency checks, probe timeouts, Istio sidecar issues, and stalled rolling deployments.
Diagnose pods stuck in Terminating — covers finalizers, force deletion, unreachable nodes, and graceful shutdown.
Resolve pods stuck in ContainerCreating — missing secrets, PVC mount failures, CNI plugin errors, and container runtime issues.
Fix EKS worker nodes that won't join — aws-auth ConfigMap, missing IAM policies, bootstrap script errors, and network issues.
Troubleshoot Ingress resources — missing controller, wrong IngressClass, empty ADDRESS, TLS issues, and backend endpoint health.
CoreDNS failures, kube-dns endpoints, NetworkPolicy blocking port 53, dnsPolicy issues, and ndots latency.
Fix inter-pod communication — Service selector/endpoint, port mismatch, NetworkPolicy, kube-proxy, and DNS.
Diagnose pods stuck in Pending — covering insufficient resources, taints, node selectors, affinity rules, and PVC binding failures.
Storage
Diagnose and fix persistent storage failures — PVC provisioning, volume mounting, and storage class configuration.
PVC stuck in Pending — StorageClass misconfiguration, no matching PV, CSI provisioner issues, and access mode mismatches.
Volume mount failures — multi-attach errors, AZ mismatches, CSI driver issues, NFS connectivity, and stale attachments.
Cluster Operations
Fix node-level, control plane, and cluster-wide operational issues.
Node NotReady — kubelet crashes, CNI plugin failures, disk/memory pressure, and control plane connectivity issues.
Pod eviction — node memory and disk pressure, missing resource limits, Priority Classes, and preventing eviction cycles.
API server unreachable — kubeconfig issues, VPN, certificate expiry, control plane downtime, and EKS private endpoint access.
Services & Networking
Fix Services that aren't reachable from inside or outside the cluster.
Scaling & Resources
Fix autoscaling, resource quotas, and deployment rollout issues.
kubectl top fails, HPA can't get metrics — Metrics Server installation, certificate errors, and kubelet connectivity.
HPA not scaling pods — unknown metrics, missing resource requests, maxReplicas limits, and stabilisation window configuration.
Quota blocking pod creation — CPU, memory, and pod count limits, right-sizing workloads, and monitoring quota usage.
Rollout not progressing — new pods crashing, ImagePullBackOff, failing readiness probes, and maxUnavailable configuration.
Guides
Practical implementation guides for real Kubernetes workflows.
Inject environment-specific values into Kubernetes manifests at deploy time — no templating engine required.
Build runbooks that actually get used during incidents — covering scope, ownership, step format, and what to leave out.
All Categories
Every article on DevOps Compass is organized into a focused category.
Pod lifecycle, workloads, probes, scheduling, and production cluster operations.
EKS, IAM, ECR, VPC architecture, and cost-aware infrastructure decisions.
Jenkins, GitLab CI, GitHub Actions, deployment strategies, and automation.
Docker, image security, registries, and container runtime debugging.
Prometheus, Grafana, Loki, alerting, and observability for production.
VPCs, IAM, TLS, network policies, and access control patterns.