Kubernetes troubleshooting
CrashLoopBackOff in Kubernetes: How to Debug
CrashLoopBackOff is not an error on its own. It means a container starts, exits, and the kubelet restarts it with a growing delay. The cause is always the process inside the container, so the work is to read why it exited.
By Muhammad Soliman, Founder at SeaGit
·
Key takeaways
- Run kubectl logs with --previous. The current container often has no useful output yet.
- The exit code narrows the cause: 1 is an application error, 126 and 127 mean a bad command, 137 is a SIGKILL (often OOM or a probe kill), and 143 is a SIGTERM.
- A container that exits with code 0 is not failing. It finished its work. A Deployment expects a foreground process that keeps running.
- A liveness probe that is too strict kills a slow-starting app, which then looks like a crash loop. Check the probe events before you change the code.
What CrashLoopBackOff means
The kubelet restarts a container whose process exits, for a pod with restartPolicy Always, which is the default for a Deployment. Each restart waits longer than the last, starting at about 10 seconds and capping at five minutes. The waiting state is CrashLoopBackOff. The restart count keeps rising while the pod is in this state.
Two states look similar and have different causes. CrashLoopBackOff means the container started and then exited. CreateContainerConfigError means the container never started, usually because a referenced secret, config map or key is missing. Read the reason before you act.
Step 1: read the logs from the last run
The current container may have just started and printed nothing. The output you need came from the run that crashed, and kubectl keeps it when you ask for --previous.
# Status, restart count and the reason for the last exit
kubectl get pod my-api-6d8f9b7c5-x2k4q -n prod
kubectl describe pod my-api-6d8f9b7c5-x2k4q -n prod | sed -n '/Containers:/,/Conditions:/p'
# Output from the run that crashed
kubectl logs my-api-6d8f9b7c5-x2k4q -n prod --previous --tail=100
# A pod with more than one container
kubectl logs my-api-6d8f9b7c5-x2k4q -n prod -c api --previousStep 2: use the exit code
Exit Code Meaning Usual cause
0 Process finished normally Not a server, or the command returned early
1 Application error Unhandled exception, bad config, failed connection at start
126 Command found but not executable Wrong file mode, or a script without a shebang
127 Command not found Wrong command or entrypoint, or a missing binary in the image
137 Killed with SIGKILL (128 + 9) OOMKilled, or a kill after the grace period ended
143 Terminated with SIGTERM (128 + 15) A liveness probe restart, or a stop requestCode 137 with Reason OOMKilled is a memory problem, and has its own guide: see the OOMKilled post. Code 137 without OOMKilled, or code 143 with a liveness probe event, points to the probe.
Step 3: match the crash to a cause
The app needs a variable or secret that is missing
The log usually shows a message such as a missing DATABASE_URL, a failed parse of a config file, or a panic on a nil value. Check that the variables exist in the running container:
kubectl exec my-api-6d8f9b7c5-x2k4q -n prod -- env | sort | head -50The dependency is not ready yet
An app that connects to a database at start and exits when the connection fails will crash-loop until the database answers. The fix is in the app: retry with a limit and a delay, or move the connection out of the startup path. An init container can wait for a port, but it should not hide a real failure.
The command or entrypoint is wrong
Exit codes 126 and 127 mean the command the pod runs does not exist in the image. Print the command the container really runs:
kubectl get pod my-api-6d8f9b7c5-x2k4q -n prod \
-o jsonpath='{.spec.containers[0].command}{" "}{.spec.containers[0].args}{"\n"}'
# Run the image locally to test the same command
docker run --rm --entrypoint sh registry.example.com/my-api:v1.4.9 -c 'which my-api'The process exits when it should keep running
A container that runs a script, then exits with code 0, is working as designed for a batch job. For a Deployment, the process must stay in the foreground. A daemon that forks and returns, or a shell command that ends, produces this loop. Run the server binary in the foreground, or use a Job for work that is meant to finish.
A liveness probe is killing the container
A liveness probe that fails restarts the container with SIGTERM or SIGKILL. The pod events show the failure:
kubectl describe pod my-api-6d8f9b7c5-x2k4q -n prod | grep -i 'probe'
kubectl get pod my-api-6d8f9b7c5-x2k4q -n prod \
-o jsonpath='{.spec.containers[0].livenessProbe}{"\n"}'For an app that takes a long time to start, add a startupProbe with a generous failureThreshold, and keep the liveness probe for the steady state. Raising initialDelaySeconds alone hides the problem for the first minute only.
Step 4: run the image yourself
A crashing container cannot be exec-ed into, so make a copy of the pod that runs a shell. kubectl debug can copy the pod spec and replace the command:
kubectl debug my-api-6d8f9b7c5-x2k4q -n prod -it \
--copy-to=my-api-debug --container=api -- sh
# Remove the copy when you are finished
kubectl delete pod my-api-debug -n prodPrevention checklist
- Log the configuration the app loaded at startup, without secret values, so a crash log explains itself.
- Make the app exit with a non-zero code and a clear message when a required setting is missing.
- Set a startupProbe for an app that needs more than a few seconds to start.
- Use a readiness probe for traffic and reserve the liveness probe for a process that is truly stuck.
How SeaGit handles this
When a deploy fails and a container is crashing, SeaGit's pod diagnosis reports the container, its exit code and the restart count. Exit code 0 gets a separate message, because it means the process finished rather than failed. The message tells you which container to look at. It does not replace the previous logs in kubectl, which hold the full output.
SeaGit does not change how your application starts. The image, the command and the environment variables come from your app configuration, so a crash loop is fixed in the app or in that configuration, not in SeaGit.