1. Introduction
A liveness probe failure in Kubernetes means your container told the control plane it was unhealthy — and Kubernetes responded by killing and restarting it. From a user perspective this looks identical to a CrashLoopBackOff: the pod keeps restarting, your application keeps interrupting. But the root cause is entirely different.
The application hasn't crashed. Kubernetes killed it because a health check returned an unexpected response, timed out, or the process it was checking wasn't ready in time. This guide walks through how liveness probes work, why they fail, and how to diagnose and fix each failure mode without breaking your deployment's ability to detect genuine application failures.
2. How Liveness Probes Work
A liveness probe is a recurring check that Kubernetes runs against a running container to decide whether it is still alive. If the probe fails failureThreshold times in a row, Kubernetes kills the container and restarts it according to the pod's restartPolicy.
Three probe mechanisms are available:
- HTTP GET — Kubernetes sends an HTTP request to a path and port. Any 2xx or 3xx status code counts as success.
- TCP Socket — Kubernetes attempts to open a TCP connection to the specified port. A successful connection = healthy.
- Exec — Kubernetes runs a command inside the container. An exit code of 0 = healthy, anything else = failure.
The key timing fields are:
| Field | What it controls | Default |
|---|---|---|
initialDelaySeconds | How long to wait after container start before running the first probe | 0 |
periodSeconds | How often to run the probe | 10 |
timeoutSeconds | How long the probe has to respond before it counts as a failure | 1 |
failureThreshold | How many consecutive failures before Kubernetes kills the container | 3 |
successThreshold | How many consecutive successes to consider the probe passing (always 1 for liveness) | 1 |
3. Common Causes
initialDelaySecondsset too low — probe fires before the application finishes starting uptimeoutSecondsset too low — application responds correctly but too slowly for the probe deadline- The health endpoint itself returns a non-2xx status when the application is under load
- The health endpoint path is wrong or has changed between deployments
- The probe port doesn't match the port the application is actually listening on
- The exec command probe exits non-zero due to a missing binary, wrong path, or transient error
- Application startup is variable — fast on warm nodes, slow on cold nodes or during image pulls
- The application is genuinely unhealthy — the probe is working as intended but the root cause needs fixing
4. Step-by-Step Diagnosis and Fix
Step 1: Confirm the probe is causing the restarts
Not every container restart is caused by a probe failure. Confirm it before changing anything:
# Check pod status and restart count
kubectl get pod <pod-name> -n <namespace>
# Describe the pod — look for probe failure events
kubectl describe pod <pod-name> -n <namespace>
# In the Events section, a liveness probe failure looks like this:
# Events:
# Warning Unhealthy 12s kubelet Liveness probe failed:
# HTTP probe failed with statuscode: 500
#
# Or for a timeout:
# Warning Unhealthy 8s kubelet Liveness probe failed:
# Get "http://10.0.0.5:8080/health": context deadline exceeded
# (Client.Timeout exceeded while awaiting headers)
#
# Followed by:
# Normal Killing 2s kubelet Container my-app failed liveness probe,
# will be restarted
If you see Liveness probe failed in the Events section and Container ... will be restarted immediately after, the probe is the kill trigger. If the Events show nothing but the pod keeps restarting, the container may be crashing for another reason — see Fix Kubernetes CrashLoopBackOff for the full diagnostic flow.
Step 2: Read the current probe configuration
Get the exact probe spec before making any changes:
# Extract the full probe config
kubectl get pod <pod-name> -n <namespace> -o jsonpath='{.spec.containers[0].livenessProbe}' | jq .
# Example output:
# {
# "httpGet": {
# "path": "/health",
# "port": 8080
# },
# "initialDelaySeconds": 5,
# "periodSeconds": 10,
# "timeoutSeconds": 1,
# "failureThreshold": 3
# }
# Also check if a startupProbe is defined — it affects liveness timing
kubectl get pod <pod-name> -n <namespace> -o jsonpath='{.spec.containers[0].startupProbe}' | jq .
Step 3: Test the probe endpoint manually
Before tuning timing values, verify the probe target is actually reachable and returns the expected response:
# For HTTP probes — exec into the container and test the endpoint directly
kubectl exec -it <pod-name> -n <namespace> -- wget -qO- http://localhost:8080/health
# or:
kubectl exec -it <pod-name> -n <namespace> -- curl -sv http://localhost:8080/health
# Check what port the app is actually listening on
kubectl exec -it <pod-name> -n <namespace> -- ss -tlnp
# or: netstat -tlnp
# For TCP socket probes — verify the port is open:
kubectl exec -it <pod-name> -n <namespace> -- nc -zv localhost 8080
# For exec probes — run the command manually:
kubectl exec -it <pod-name> -n <namespace> -- /bin/sh -c "your-probe-command; echo exit:$?"
If the endpoint returns a non-2xx status, is unreachable, or the command exits non-zero, you've found the issue. The fix may be in the application health logic, not in the probe configuration.
Step 4: Fix — Increase initialDelaySeconds or add a startupProbe
The most common liveness probe failure in new deployments is a probe that fires before the application finishes starting. If the Events show failures within the first 30 seconds of pod start, this is almost certainly the cause:
# Option A: increase initialDelaySeconds (simpler, less precise)
livenessProbe:
httpGet:
path: /health
port: 8080
initialDelaySeconds: 60 # Give the app 60s to start before first probe
periodSeconds: 15
timeoutSeconds: 3
failureThreshold: 3
# Option B: use a startupProbe (recommended for variable startup times)
# The startupProbe disables the livenessProbe until it succeeds.
# This allows slow-starting apps (JVM, apps loading large configs) to
# take as long as they need without being killed by liveness.
startupProbe:
httpGet:
path: /health
port: 8080
failureThreshold: 30 # Allow up to 30 * periodSeconds = 5 minutes to start
periodSeconds: 10 # Check every 10 seconds during startup
livenessProbe:
httpGet:
path: /health
port: 8080
initialDelaySeconds: 0 # startupProbe guards this — no delay needed
periodSeconds: 15
timeoutSeconds: 3
failureThreshold: 3
Step 5: Fix — Increase timeoutSeconds for slow health endpoints
If the Events show context deadline exceeded or i/o timeout, the probe is timing out rather than receiving an explicit failure. The default timeoutSeconds: 1 is very tight for any endpoint that does real work:
livenessProbe:
httpGet:
path: /health
port: 8080
initialDelaySeconds: 30
periodSeconds: 15
timeoutSeconds: 5 # Give the endpoint 5 seconds to respond
failureThreshold: 3
# Important: a /health endpoint that queries a database or external
# service to determine health will be slow under load. Consider making
# /health a lightweight check (process is alive, can accept connections)
# and moving dependency checks to /ready (readiness probe).
Step 6: Fix — Correct a wrong path or port
If the probe returns 404 or connection refused, the path or port in the probe spec is wrong:
# Find the correct port the application listens on:
kubectl get pod <pod-name> -n <namespace> -o jsonpath='{.spec.containers[0].ports}'
# Check what paths the application exposes (for HTTP probes):
kubectl exec -it <pod-name> -n <namespace> -- curl -s http://localhost:8080/ # Try the root path
kubectl exec -it <pod-name> -n <namespace> -- curl -s http://localhost:8080/healthz # Common alternatives
kubectl exec -it <pod-name> -n <namespace> -- curl -s http://localhost:8080/actuator/health # Spring Boot
# Update the probe in your Deployment manifest:
livenessProbe:
httpGet:
path: /actuator/health # corrected path
port: 8080 # verified port
Step 7: Fix — Handle genuinely unhealthy applications
If the probe is correctly configured and the application is still returning failures, the application itself is in a bad state. Common causes include:
- The app is stuck in an unrecoverable state (deadlock, connection pool exhausted) and the health endpoint reflects this correctly
- A dependency (database, cache, external API) is down and the health endpoint incorrectly reports this as the application being unhealthy
- The app is genuinely crashing — check
kubectl logs <pod-name> --previousfor errors before the restart
# Get logs from the container just before the last restart
kubectl logs <pod-name> -n <namespace> --previous
# Watch live logs to see what happens at the moment of probe failure
kubectl logs <pod-name> -n <namespace> -f
# Check for application-level errors at the time of each restart:
kubectl get events -n <namespace> --sort-by='.lastTimestamp' | tail -20
5. Verification Steps
After applying your fix, verify the probe stabilises:
# Watch for probe failure events — there should be none
kubectl get events -n <namespace> --field-selector reason=Unhealthy -w
# Watch pod restarts stop incrementing
kubectl get pod <pod-name> -n <namespace> -w
# Expected: RESTARTS count freezes, STATUS stays Running
# Verify the updated probe configuration is live
kubectl get pod <pod-name> -n <namespace> -o jsonpath='{.spec.containers[0].livenessProbe}' | jq .
# Manually trigger the probe target to confirm it responds correctly
kubectl exec -it <pod-name> -n <namespace> -- curl -sv http://localhost:8080/health
# Expected: HTTP 200 response within your timeoutSeconds value
6. Common Mistakes
- Setting
initialDelaySecondsto a very large value (120+) as a shortcut — this delays genuine failure detection long after startup. UsestartupProbeinstead. - Making the liveness probe the same as the readiness probe — they serve different purposes. Liveness failure kills the container; readiness failure removes it from traffic. The thresholds and targets should usually differ.
- Putting downstream dependency checks in the liveness endpoint — when a dependency fails, your healthy pods get killed unnecessarily, turning a partial outage into a full one.
- Setting
failureThreshold: 1— a single slow response or transient hiccup will immediately kill the container. Use at least 3. - Not testing the probe manually before deploying — probes are invisible until they fail in production. Test the endpoint with
curlorwgetfrom inside the container first. - Confusing liveness and startup probe failures — if you have a
startupProbeand the container is killed within the first few minutes, checkkubectl describe podcarefully to see whether it's the startup or liveness probe that failed.
7. Prevention Tips
- Add a
/healthor/healthzendpoint to every application from day one — it should return 200 if the process is alive and ready to handle requests, nothing more - Use
startupProbefor any application with variable startup time (JVM apps, services loading large configs, ML models at init) - Keep liveness probes lightweight — the endpoint should respond in under 100ms under normal load and should never query external dependencies
- Separate liveness from readiness: liveness checks whether the app is alive; readiness checks whether it can serve traffic right now
- Set
timeoutSecondsto at least 3–5 seconds for any HTTP probe — the default of 1 second is too tight for any real-world service under load - Monitor probe failure events with Prometheus: the metric
kube_pod_container_status_restarts_totalcombined withkube_eventsforUnhealthyreason gives early warning before CrashLoopBackOff sets in - Test probes in staging with realistic load before enabling them in production — a probe that passes under no-load can fail under the latency of real traffic
8. Summary
Liveness probe failures cause Kubernetes to kill and restart your container. The fix is almost never "disable the probe" — it's either tuning the timing so the probe doesn't fire too early or too fast, fixing the health endpoint logic, or correcting a wrong path or port.
| Symptom / Error | Most likely cause | Fix |
|---|---|---|
| Restarts in first 30–60s of pod life | initialDelaySeconds too low |
Increase delay or add startupProbe |
context deadline exceeded |
timeoutSeconds too low |
Increase to 3–5s; make health endpoint faster |
| HTTP 404 / connection refused | Wrong path or port in probe spec | Verify port with ss -tlnp; test path with curl |
| HTTP 503 when dependency is down | Dependency check in liveness endpoint | Move dependency checks to readiness probe only |
| Probe passes but restarts continue | Application crashing for another reason | Check kubectl logs --previous; see CrashLoopBackOff guide |
Start with kubectl describe pod and read the Events section. The probe failure message will tell you exactly what the probe tried and what response it got.