1. Introduction
A NotReady node is a cluster-level emergency. Any pods currently running on that node may be unreachable or terminated. New pods won't be scheduled to it. If enough nodes go NotReady simultaneously, your cluster loses capacity and workloads fail entirely. The cause can be a failed kubelet, a broken network plugin, resource exhaustion, or loss of connectivity to the control plane.
This guide walks through the complete diagnostic path: from what kubectl get nodes tells you, through SSH-level investigation, to resolution and pod rescheduling.
2. What NotReady Actually Means
A node is Ready when the kubelet reports all of the following: sufficient disk, memory, and PID resources; the container runtime is responsive; the network plugin is configured; and the node can communicate with the control plane. If any of these conditions fail for more than node-monitor-grace-period (default 40s), the node transitions to NotReady.
3. Common Causes
- kubelet process stopped or crashed on the node
- CNI plugin failure — the network plugin can't configure pod networking
- Node has MemoryPressure, DiskPressure, or PIDPressure
- Container runtime (containerd, Docker) is unresponsive
- Node lost network connectivity to the control plane API server
- Certificate expiry — kubelet TLS certificates expired
- Clock skew between node and control plane exceeding tolerance
- Kernel panic or hardware failure
- Full disk preventing the kubelet from writing state
4. Step-by-Step Diagnosis and Fix
Step 1: Read the node conditions
# Check all node conditions
kubectl describe node <node-name> | grep -A 20 "Conditions:"
# Look for:
# MemoryPressure: False (True = low memory)
# DiskPressure: False (True = low disk)
# PIDPressure: False (True = too many processes)
# Ready: False (this is the main signal)
# NetworkUnavailable: False (True = CNI failure)
# Also check Events at the bottom:
kubectl describe node <node-name> | tail -20
Step 2: Check the kubelet on the node
# SSH into the node (or use SSM Session Manager for AWS)
ssh -i key.pem ec2-user@<node-ip>
# Check kubelet service status
systemctl status kubelet
journalctl -u kubelet -n 100 --no-pager
# Common kubelet errors:
# "failed to get node info" = API server unreachable
# "certificate has expired" = TLS cert issue
# "runtime not responding" = containerd/Docker down
# "node disk pressure" = disk full
# Restart kubelet if it's stopped
systemctl restart kubelet
systemctl enable kubelet
Step 3: Check container runtime
# Check containerd
systemctl status containerd
journalctl -u containerd -n 50 --no-pager
# Test if containerd responds
crictl info
# Check Docker (if used)
systemctl status docker
docker info
# Restart if needed
systemctl restart containerd
# Or: systemctl restart docker
Step 4: Check CNI plugin
# From the control plane: check CNI DaemonSet pods on the affected node
NODE=<node-name>
kubectl get pods -n kube-system --field-selector spec.nodeName=$NODE | grep -E "calico|flannel|aws-node|cilium|weave"
# Check CNI pod logs
kubectl logs -n kube-system <cni-pod> --tail=50
# On the node itself: check CNI config files
ls /etc/cni/net.d/
cat /etc/cni/net.d/*.conf
# Restart CNI pod to force re-initialization:
kubectl delete pod -n kube-system <cni-pod-on-node>
Step 5: Check control plane connectivity from the node
# On the node, test connectivity to the API server
# Get API server endpoint:
kubectl cluster-info 2>/dev/null | grep "Kubernetes master\|Kubernetes control plane"
# Test from the node:
curl -k https://<api-server-endpoint>:6443/healthz
# Expected: ok
# Check for network issues to the API server:
traceroute <api-server-ip>
# For EKS: verify the node can reach the VPC endpoint
nslookup <eks-cluster-endpoint>
Step 6: Check disk and memory
# On the node:
df -h # disk usage — look for 100% on any filesystem
free -m # memory
df -i # inode usage
# Find what's using disk:
du -sh /var/lib/docker/* 2>/dev/null | sort -rh | head -10
du -sh /var/log/containers/* | sort -rh | head -10
# Emergency disk cleanup:
docker system prune -af 2>/dev/null || crictl rmi --prune 2>/dev/null
journalctl --vacuum-size=200M
Step 7: Drain, fix, and return node to service
# Cordon to prevent new pods scheduling during repair
kubectl cordon <node-name>
# Drain existing pods (if node will be rebooted)
kubectl drain <node-name> --ignore-daemonsets --delete-emptydir-data --force
# After fixing: uncordon to allow scheduling again
kubectl uncordon <node-name>
# Watch node recover
kubectl get nodes -w
5. Verification Steps
# Node should show Ready
kubectl get nodes
# NAME STATUS ROLES AGE VERSION
# node-1 Ready <none> 10d v1.29.x
# All node conditions should be False except Ready
kubectl describe node <node-name> | grep -A 8 "Conditions:"
# Pods should reschedule and become Running
kubectl get pods -A --field-selector spec.nodeName=<node-name>
6. Common Mistakes
- Rebooting the node without diagnosing the root cause — the same issue will recur after reboot
- Not draining the node before rebooting — pods don't get graceful shutdown
- Assuming NotReady means the node is gone — often the kubelet just needs a restart
- Not checking CNI plugin pods on the node — a failed CNI pod causes NetworkUnavailable without obvious kubelet errors
- Ignoring clock skew — even a 10-minute difference between node and control plane can cause certificate validation failures
7. Prevention Tips
- Alert on node conditions (
MemoryPressure,DiskPressure) before they cause NotReady transitions - Set up node problem detector to surface kernel and runtime errors before the kubelet fails
- Configure log rotation and image garbage collection to prevent disk pressure
- Use managed node groups (EKS, GKE, AKS) where the provider handles node recovery automatically
- Monitor API server connectivity from nodes — a network partition between nodes and the control plane is a common cause of mass NotReady
- Check Fix Kubernetes API Server Connection Refused if control plane connectivity is the root cause
8. FAQ
Multiple nodes went NotReady at the same time. What happened?
Simultaneous NotReady across multiple nodes almost always indicates: (1) a control plane failure (API server, etcd), (2) a network partition between the nodes and control plane, or (3) a cluster-wide issue like a certificate expiry. Check the control plane components first: kubectl get pods -n kube-system. If the control plane is unreachable, SSH directly to a node and check journalctl -u kubelet.
The node shows Ready again but pods aren't rescheduling. Why?
Pods from a recovered NotReady node are rescheduled by the controller manager, but only if the pods have a managing controller (Deployment, ReplicaSet, etc.). Standalone pods are not automatically rescheduled. Also check that the node isn't cordoned (kubectl get nodes shows SchedulingDisabled) — run kubectl uncordon <node> to allow scheduling.
9. Summary
| Condition showing | Root cause | Fix |
|---|---|---|
| NetworkUnavailable=True | CNI plugin failure | Restart CNI DaemonSet pod on node |
| MemoryPressure=True | Node out of memory | Evict pods; add capacity; right-size workloads |
| DiskPressure=True | Node disk full | Clean logs/images; enable log rotation |
| kubelet stopped | Process crash | systemctl restart kubelet on the node |
| API server unreachable | Network or control plane issue | Check VPC routes, SGs, control plane health |
Explore More in This Category
Explore more in this category: Kubernetes guides. Browse all DevOps Compass articles or jump to: Kubernetes, AWS, CI/CD, Containers, Monitoring, Networking.