1. Introduction
An OOMKilled pod in Kubernetes means the Linux kernel ran out of patience with your container. It hit the memory limit you configured, the kernel's OOM killer fired, and the process was terminated with exit code 137. Kubernetes then restarts the container — and if the underlying cause isn't fixed, it enters CrashLoopBackOff, restarting over and over.
OOMKilled can be caused by three fundamentally different things: a memory limit that was set too low for the workload, a genuine memory leak in the application, or a JVM (or other runtime) that sizes itself to the node rather than respecting the container's limits. Each requires a different fix. This guide walks through how to identify which one you're dealing with and how to resolve it.
2. What OOMKilled Actually Means
In Kubernetes, every container can have a resources.limits.memory value in its pod spec. When a container's memory usage reaches that limit, the Linux kernel's Out-Of-Memory (OOM) killer terminates the container's process. The exit code is always 137 (128 + signal 9, SIGKILL).
Kubernetes records this as the reason OOMKilled in the pod's last state. Unlike a regular application crash, OOMKilled is not something your application can catch or handle — the kernel kills it unconditionally, without waiting for a graceful shutdown.
There are two distinct OOM scenarios in Kubernetes, and they produce different behaviour:
- Container-level OOM: The container hits its own
limits.memory. Only that container is killed. This is the most common scenario. - Node-level OOM: The node itself runs out of memory. The kernel picks a process to kill cluster-wide. This is rarer and usually means your resource requests are set too low, allowing over-commitment.
3. Common Causes
- Memory limit set too low — the workload legitimately needs more than was allocated
- Memory leak in application code — usage grows unbounded until the limit is hit
- JVM auto-sizing to node memory rather than container limits (pre-Java 10 behaviour)
- No memory limit set at all — pod competes with all other workloads on the node
- Memory requests set too low — the scheduler places pods on nodes that can't actually support their peak usage
- Large in-memory caches or buffers that aren't bounded by configuration
- Sidecar containers (logging agents, service mesh proxies) accumulating memory and collectively pushing the pod over the limit
4. Step-by-Step Diagnosis and Fix
Step 1: Confirm OOMKilled is the cause
First, confirm you're actually dealing with an OOMKilled event and not another type of crash:
# Check pod status and restart count
kubectl get pod <pod-name> -n <namespace>
# Look at the last terminated state — this is the key command
kubectl describe pod <pod-name> -n <namespace>
# In the output, look for this in the container status section:
# Last State: Terminated
# Reason: OOMKilled
# Exit Code: 137
# Started: Mon, 21 Apr 2025 09:14:02 +0000
# Finished: Mon, 21 Apr 2025 09:14:45 +0000
Exit code 137 with reason OOMKilled confirms a memory kill. If the exit code is something else — 1, 2, or 127 — you're likely dealing with a different issue. See the CrashLoopBackOff guide for a full exit code reference.
Step 2: Check the current memory limit and actual usage
Before changing anything, understand what limit is set and whether actual usage justifies it:
# Check the memory limit set on the container
kubectl get pod <pod-name> -n <namespace> \
-o jsonpath='{.spec.containers[0].resources}'
# Example output:
# {
# "limits": {"memory": "256Mi"},
# "requests": {"memory": "128Mi"}
# }
# Check live memory usage with metrics-server
kubectl top pod <pod-name> -n <namespace>
# Check across all pods in the namespace
kubectl top pods -n <namespace> --sort-by=memory
If kubectl top shows the pod was consistently using 240–250Mi before getting killed, and the limit is 256Mi, the limit is probably just slightly too low. If it shows usage growing steadily over hours and approaching the limit at the point of each kill, you likely have a memory leak.
Step 3: Look at memory usage over time
A single kubectl top snapshot doesn't tell you whether memory usage is stable or growing. For a clearer picture, check Prometheus and Grafana if you have them, or use the following approach to observe the trend:
# Watch memory usage in real time (requires metrics-server)
watch -n 5 "kubectl top pod <pod-name> -n <namespace>"
# In Prometheus, the relevant metric is:
# container_memory_working_set_bytes — excludes reclaimable cache,
# closer to what Kubernetes uses for OOM decisions
#
# Useful PromQL query to spot leaks:
# container_memory_working_set_bytes{
# namespace="<namespace>",
# pod=~"<pod-prefix>.*",
# container="<container>"
# }
A flat or slightly fluctuating line means the limit is too low for a stable workload. A consistently upward-sloping line means a memory leak — raising the limit will only delay the next kill, not prevent it.
Step 4: Fix — Increase the memory limit (if the workload is stable)
If the usage pattern is stable (not growing) and the limit is too tight, increase it in your Deployment or pod spec:
# In your Deployment manifest:
spec:
containers:
- name: my-app
resources:
requests:
memory: "256Mi" # What the scheduler uses for placement
limits:
memory: "512Mi" # The hard ceiling the kernel enforces
# Apply the change
kubectl apply -f deployment.yaml
# Or patch it directly (use sparingly — prefer manifest changes in git):
kubectl set resources deployment <name> -n <namespace> \
--limits=memory=512Mi \
--requests=memory=256Mi
Step 5: Fix — Java / JVM memory issues
Java is the most common source of OOMKilled that isn't a "real" memory leak. Prior to Java 10, the JVM detected the node's total memory and sized the heap accordingly — completely ignoring the container's limits.memory. On a node with 32 GiB of RAM, a JVM in a container limited to 512Mi would attempt to allocate multiple gigabytes of heap.
Java 10+ supports container-aware sizing via -XX:+UseContainerSupport (enabled by default). If you're running an older JVM, or if the auto-sizing is still too aggressive, set the heap explicitly:
# Recommended: set max heap to ~75% of your container memory limit
# For a 512Mi limit, set -Xmx384m
env:
- name: JAVA_OPTS
value: "-Xms128m -Xmx384m -XX:+UseContainerSupport"
# Or using the JVM container support flags directly:
# -XX:MaxRAMPercentage=75.0 (percentage of container memory)
# -XX:InitialRAMPercentage=25.0
# Verify the JVM sees the right memory inside a running container:
kubectl exec -it <pod-name> -n <namespace> -- \
java -XX:+PrintFlagsFinal -version 2>&1 | grep -i maxheapsize
Step 6: Fix — Investigate and address a memory leak
If usage is consistently growing regardless of limit size, you have a memory leak. The correct fix is in the application code, not the Kubernetes configuration. However, you can take steps to diagnose it without access to a production heap dump:
# Trigger a heap dump from a running JVM container (Java)
kubectl exec -it <pod-name> -n <namespace> -- \
jmap -dump:format=b,file=/tmp/heap.hprof <jvm-pid>
# Copy the dump out for analysis
kubectl cp <namespace>/<pod-name>:/tmp/heap.hprof ./heap.hprof
# For Go applications, trigger a pprof memory profile:
kubectl exec -it <pod-name> -n <namespace> -- \
curl -s http://localhost:6060/debug/pprof/heap > /tmp/heap.prof
# For Node.js, use --inspect and connect remotely,
# or add a /heap-snapshot endpoint using v8.writeHeapSnapshot()
While the real fix is being developed, a common short-term mitigation is to configure a memory limit high enough to give the leak time to be identified, combined with a liveness probe that triggers a restart before the container hits the limit:
livenessProbe:
httpGet:
path: /healthz
port: 8080
initialDelaySeconds: 30
periodSeconds: 15
failureThreshold: 3
# A well-designed /healthz endpoint can return 503 when memory
# usage exceeds a threshold, triggering a graceful restart
# before the kernel kills the process.
Step 7: Right-size using VPA recommendations
If you don't have historical usage data or you're setting limits for the first time, the Kubernetes Vertical Pod Autoscaler (VPA) in recommendation mode gives you data-driven sizing without automatically changing anything:
# Install VPA (if not already installed)
# https://github.com/kubernetes/autoscaler/tree/master/vertical-pod-autoscaler
# Create a VPA object in recommendation mode (does not modify pods)
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
name: my-app-vpa
namespace: <namespace>
spec:
targetRef:
apiVersion: apps/v1
kind: Deployment
name: my-app
updatePolicy:
updateMode: "Off" # Recommendation only — does not auto-apply
# After a few hours/days of data, check the recommendations:
kubectl describe vpa my-app-vpa -n <namespace>
# Example output:
# Recommendation:
# Container Recommendations:
# Container Name: my-app
# Target:
# Memory: 420Mi <-- use this as your limit
# Lower Bound:
# Memory: 240Mi
# Upper Bound:
# Memory: 680Mi
5. Verification Steps
After applying your fix, confirm the pod stays stable:
# Watch the pod — restarts should stop incrementing
kubectl get pod <pod-name> -n <namespace> -w
# Expected: RESTARTS count stops growing, STATUS stays Running
# NAME READY STATUS RESTARTS AGE
# myapp-abc 1/1 Running 3 48m <-- restarts frozen
# Verify the new limits are applied
kubectl get pod <pod-name> -n <namespace> \
-o jsonpath='{.spec.containers[0].resources}'
# Monitor memory usage over the next 30–60 minutes
kubectl top pod <pod-name> -n <namespace>
# Check there are no new OOMKilled events
kubectl describe pod <pod-name> -n <namespace> | grep -A 5 "Last State"
# Expected: Last State section should be empty or show a prior Terminated
# with no new OOMKilled entry
6. Common Mistakes
- Increasing the memory limit without first determining whether the issue is a leak or an under-set limit — you may just be delaying the next OOMKill
- Using
kubectl set resourcesdirectly on a running pod instead of updating the Deployment manifest — the change is overwritten on the next rollout - Setting
limits.memorywithout settingrequests.memory— Kubernetes treats the limit as the request when no request is specified, which can prevent the pod from scheduling on nodes that actually have capacity - Forgetting sidecar containers in the limit calculation — Istio/Envoy proxies, Datadog agents, and Fluent Bit all consume memory that counts against the pod total
- Not accounting for JVM non-heap memory — setting
-Xmxto equal the container limit always causes OOMKilled because the total JVM footprint exceeds the heap size - Assuming a stable workload has a leak just because it grows after a restart — some applications build caches on startup that plateau after a warmup period
7. Prevention Tips
- Set both
requestsandlimitson every container — pods with no requests are scheduled unpredictably, and pods with no limits can consume unlimited node memory - Use VPA in recommendation mode to get data-driven baseline sizes before setting limits manually
- Add a Prometheus alert on
container_memory_working_set_bytesthat fires when a container hits 85% of its memory limit — gives you time to act before the OOM kill - For Java workloads, always set
-XX:MaxRAMPercentageexplicitly and verify with-XX:+PrintFlagsFinal— never rely on auto-detection alone - Apply a runbook for recurring OOMKilled events — document the typical memory usage patterns for each service so on-call engineers know whether a given kill is expected or anomalous
- Use namespace-level
LimitRangeobjects to enforce default requests and limits, ensuring no pod in the namespace goes without them - Profile memory usage in staging under realistic load before setting production limits — synthetic benchmarks rarely represent actual peak consumption
8. Summary
OOMKilled always means the container exceeded its memory limit and the kernel terminated it. The diagnostic question to answer first is: is the limit too low, or is memory genuinely growing unbounded?
| Cause | Signal | Fix |
|---|---|---|
| Limit too low | Flat/stable memory usage at time of kill | Increase limits.memory; set requests proportionally |
| Memory leak | Memory grows steadily before each kill | Profile and fix in application; use liveness probe as interim mitigation |
| JVM over-allocation | Java app, large heap, old JVM version | Set -XX:MaxRAMPercentage=75 or explicit -Xmx; leave 25% for non-heap |
| Sidecar pressure | Multiple containers, combined usage over limit | Add per-container limits; audit sidecar memory consumption |
| No limit set | Node-level OOM, other pods affected | Add limits.memory to all containers; use LimitRange as a safety net |
Start with kubectl describe pod to confirm exit code 137 and reason OOMKilled, then check kubectl top pod to understand the usage pattern before reaching for a higher limit.