1. Introduction
CrashLoopBackOff is one of the most common Kubernetes error states — and one of the most misunderstood. Unlike ImagePullBackOff, the container actually starts. It just crashes, over and over again, and Kubernetes keeps restarting it with increasing delays.
The frustrating part is that CrashLoopBackOff itself tells you almost nothing. The real cause is buried in the container logs, exit codes, or pod events. This guide gives you a systematic way to find it.
We'll cover every common root cause with the exact commands to diagnose it and the steps to fix it.
2. What CrashLoopBackOff Actually Means
When a pod enters CrashLoopBackOff, it means:
- The container started successfully (image pulled, container runtime launched it)
- The container exited with a non-zero exit code or was killed by the system
- Kubernetes attempted to restart it — and it crashed again
- Kubernetes is now backing off restarts using an exponential delay (10s, 20s, 40s... up to 5 minutes)
The backoff is Kubernetes protecting cluster resources from a tight crash loop consuming CPU and memory. The pod stays visible and scheduled — it's just being restarted less frequently as the restarts accumulate.
3. Common Causes
- Application crashes on startup due to a missing or invalid environment variable
- A required secret or ConfigMap is missing or has a wrong key
- Incorrect command or entrypoint in the container spec
- Container runs to completion — it's a Job task in a Deployment by mistake
- OOMKilled — the container exceeds its memory limit and is killed by the kernel
- Liveness probe misconfigured — killing a healthy container too early
- Dependency not ready on startup — app fails because the database isn't reachable yet
- Permissions issue — app can't read a mounted volume or write to a path
- Port binding failure — app tries to bind to a port already in use
4. Step-by-Step Diagnosis and Fix
Step 1: Confirm the state and check restart count
Start by getting a quick picture of what's happening:
kubectl get pod <pod-name> -n <namespace>
# Example output:
# NAME READY STATUS RESTARTS AGE
# myapp-7f9bc 0/1 CrashLoopBackOff 8 14m
# High restart count = it's been crashing repeatedly
# Check all pods across namespaces if you're unsure which one: kubectl get pods -A | grep CrashLoop
Step 2: Read the container logs
This is the most important step. The logs from the last crash attempt almost always reveal the cause:
# Logs from the current (crashing) container kubectl logs <pod-name> -n <namespace>
# Logs from the PREVIOUS crashed container — often more useful kubectl logs <pod-name> -n <namespace> --previous
# If there are multiple containers in the pod: kubectl logs <pod-name> -n <namespace> -c <container-name> --previous
Read the last few lines carefully. Look for:
- Panic messages or stack traces (application crash)
- "No such file or directory" (wrong path, missing binary, wrong entrypoint)
- "Permission denied" (volume mount or filesystem permission issue)
- "Connection refused" or timeout errors (dependency not ready)
- "Error loading config" or "env var not set" (missing configuration)
Step 3: Check the exit code
The exit code tells you how the container terminated. Get it from the pod description:
kubectl describe pod <pod-name> -n <namespace>
# Look for the Last State section:
# Last State: Terminated
# Reason: Error
# Exit Code: 1
# Started: Mon, 21 Apr 2025 10:14:02 +0000
# Finished: Mon, 21 Apr 2025 10:14:03 +0000
| Exit Code | What it means |
|---|---|
| 0 | Container exited cleanly — it completed and stopped. Wrong workload type (use a Job, not Deployment). |
| 1 | General application error. Read the logs — usually a config issue, unhandled exception, or startup failure. |
| 2 | Misuse of shell built-in or script error. Check your entrypoint/command for shell syntax issues. |
| 126 | Command found but not executable. Permissions issue on the binary or script. |
| 127 | Command not found. Wrong entrypoint path, missing binary in image, or typo in command. |
| 137 | SIGKILL — process was killed. Usually OOMKilled (memory limit exceeded) or a manual kill. |
| 139 | Segmentation fault. Application-level crash — typically a bug in native code. |
| 143 | SIGTERM — container was gracefully terminated but didn't shut down in time. |
Step 4: Check the entrypoint and command
A common cause of fast crashes (exit code 127 or 2) is a wrong command or entrypoint in the pod spec. Verify what the container is actually trying to run:
kubectl get pod <pod-name> -n <namespace> -o jsonpath=\ '{.spec.containers[0].command} {.spec.containers[0].args}'
# Also check what the image's default entrypoint is: docker inspect <image>:<tag> --format='{{.Config.Entrypoint}} {{.Config.Cmd}}'
If your pod spec overrides the command, make sure the binary exists in the image at that exact path. If you're running a shell script, ensure it's executable and uses the correct interpreter line.
Step 5: Check for missing env vars and secrets
Many apps crash immediately on startup if a required environment variable is missing or empty. If your app uses AWS APIs, missing IRSA credentials are a common cause — see How to Use IRSA in EKS for the correct setup. Check what the pod has access to:
# List env vars configured on the container kubectl get pod <pod-name> -n <namespace> -o jsonpath=\ '{.spec.containers[0].env}'
# Check that referenced secrets exist kubectl get secret <secret-name> -n <namespace>
# Check that referenced ConfigMaps exist kubectl get configmap <configmap-name> -n <namespace>
# Verify a specific secret has the expected key kubectl get secret <secret-name> -n <namespace> -o jsonpath='{.data}' | jq 'keys'
If a secret or ConfigMap key is missing, the pod spec using valueFrom will cause the container to fail at injection time — and the logs may be sparse. The pod Events (from kubectl describe) will show a clearer error in this case.
Step 6: Check for OOMKilled
If the exit code is 137 and the reason in kubectl describe shows OOMKilled, the container exceeded its memory limit and was killed by the Linux OOM killer:
kubectl describe pod <pod-name> -n <namespace>
# Look for:
# Last State: Terminated
# Reason: OOMKilled
# Exit Code: 137
# Check the current memory limit: kubectl get pod <pod-name> -n <namespace> -o jsonpath=\ '{.spec.containers[0].resources}'
Fixes for OOMKilled:
- Increase the memory limit in the pod spec if the current limit is too low for the workload
- Profile the application to identify a memory leak if usage grows unbounded
- Set JVM heap size explicitly if running Java — the JVM may auto-size to node memory, not container limits
Step 7: Check liveness probe configuration
A misconfigured liveness probe can put a healthy app into CrashLoopBackOff by killing the container before it finishes starting. Check the probe settings:
kubectl get pod <pod-name> -n <namespace> -o jsonpath=\ '{.spec.containers[0].livenessProbe}'
# Example output showing a probe that fires too early:
# {"httpGet":{"path":"/health","port":8080},
# "initialDelaySeconds":3,
# "periodSeconds":5,
# "failureThreshold":3}
If initialDelaySeconds is too short for your app's startup time, the probe fires before the app is ready, fails, and Kubernetes kills the container. Fix by increasing initialDelaySeconds, or better — use a separate startupProbe that disables the liveness probe until the container has had time to initialise:
startupProbe: httpGet: path: /health port: 8080 failureThreshold: 30 # allow up to 30 x periodSeconds to start periodSeconds: 10 livenessProbe: httpGet: path: /health port: 8080 initialDelaySeconds: 0 # startupProbe guards this now periodSeconds: 15 failureThreshold: 3
Step 8: Check for exit code 0 (completed container)
If the exit code is 0 and the pod keeps restarting, the container is completing successfully — but it's in a Deployment, which expects containers to run continuously. Kubernetes sees it exit and restarts it.
This happens when a one-off script, migration job, or batch task is accidentally deployed as a Deployment. The fix is to use the correct workload type:
- Use a Job for one-off tasks that should run to completion
- Use a CronJob for tasks that run on a schedule
- Use a Deployment only for long-running processes that should always be up
5. Verification Steps
After applying your fix, confirm the pod stabilises:
# Watch the pod recover in real time kubectl get pod <pod-name> -n <namespace> -w
# Expected recovery sequence:
# NAME READY STATUS RESTARTS
# myapp-7f9bc 0/1 CrashLoopBackOff 9
# myapp-7f9bc 0/1 Error 9
# myapp-7f9bc 0/1 Running 9
# myapp-7f9bc 1/1 Running 9 <-- stable
# Confirm it stays Running without incrementing restarts: kubectl get pod <pod-name> -n <namespace>
# For a Deployment, check rollout status: kubectl rollout status deployment/<name> -n <namespace>
6. Common Mistakes
- Reading only the current logs — always use --previous if the container has already crashed at least once
- Ignoring the exit code — it narrows down the cause significantly before you look at anything else
- Increasing memory limits without profiling first — you may be masking a memory leak
- Setting initialDelaySeconds very high as a workaround for slow startup — use startupProbe instead
- Fixing the wrong pod — if you're using a Deployment, changes to the pod directly don't persist; update the Deployment spec
- Assuming the issue is in Kubernetes — CrashLoopBackOff is almost always an application-level problem
7. Prevention Tips
- Add a /health or /ready endpoint to your application and configure proper probes from day one
- Set resource requests and limits on every container — this prevents OOMKilled and gives the scheduler accurate information
- Use startupProbe for any application with variable startup time (JVM, apps loading large configs on boot)
- Validate that all required secrets and ConfigMaps exist before deploying — use pre-deployment checks in your CI/CD pipeline
- Run kubectl apply --dry-run=server before deploying to catch spec errors early
- Add structured logging to your application so startup failures produce readable output, not just a non-zero exit — and consider capturing the fix in a runbook so the next on-call engineer has a clear path to follow
- Use init containers for dependency checks — an init container that waits for a database to be ready prevents the main container from crashing on startup
8. Summary
CrashLoopBackOff means your container keeps exiting. The fix depends entirely on why it's exiting — which you find by reading the logs and the exit code. Here's the quick reference:
| Cause | Signal to look for | Fix |
|---|---|---|
| Exit code 0 | Status: Completed | Use a Job, not a Deployment |
| Exit code 1 | App error in logs | Fix config, env vars, or app bug |
| Exit code 127 | Command not found | Fix entrypoint path in pod spec |
| Exit code 137 | Reason: OOMKilled | Increase memory limit or fix leak |
| Probe failure | Events: Liveness probe failed | Tune initialDelaySeconds or add startupProbe |
| Missing secret | Events: secret not found | Create the missing secret/ConfigMap key |
| Dependency | Timeout errors in logs | Use init containers for readiness wait |
Start with kubectl logs --previous and kubectl describe pod. Between those two commands, you'll find the cause of the vast majority of CrashLoopBackOff issues.
Explore More in This Category
Explore more in this category: Kubernetes guides. Browse all DevOps Compass articles or jump to a related area: Kubernetes, AWS, CI/CD, Containers, Monitoring, Networking.