Kubernetes troubleshooting
ImagePullBackOff in Kubernetes: Find and Fix It
ImagePullBackOff means the kubelet could not pull a container image and is retrying with a growing delay. The container never started, so the fault is in the image reference, the registry credentials or the network path to the registry. The application code is not the cause.
By Muhammad Soliman, Founder at SeaGit
·
Key takeaways
- ErrImagePull is the first failure. ImagePullBackOff is the same failure after the kubelet starts backing off between retries.
- The event text tells you which cause you have: NotFound means the tag is wrong, "pull access denied" or 401 means credentials, and i/o timeout or no such host means the network.
- On EKS, pods pull from ECR with the node IAM role, so an ECR image in the same account needs no secret. A registry outside ECR needs an imagePullSecret.
- A manually created ECR docker-registry secret expires after 12 hours. Do not use one for a long-running workload.
- Pin an immutable tag such as a git SHA. A :latest tag hides which image is running.
Read the pull error first
The reason is in the pod events, not the pod status. Start with the image the pod is trying to pull, then read the last events for that pod.
kubectl get pod my-api-7c9f -n prod \
-o jsonpath='{.spec.containers[*].image}{"\n"}'
kubectl describe pod my-api-7c9f -n prod | sed -n '/Events:/,$p'
kubectl get events -n prod \
--field-selector involvedObject.name=my-api-7c9f \
--sort-by=.lastTimestamp | tail -20Match the message to a cause in the table below. The text is illustrative; copy the exact line from your own events.
# Tag does not exist in the registry
Failed to pull image "111122223333.dkr.ecr.eu-west-1.amazonaws.com/my-api:v1.4.9": ... not found
# Credentials missing or rejected
pull access denied, repository does not exist or may require 'docker login'
failed to authorize: ... 401 Unauthorized
# Network path to the registry is broken
dial tcp: lookup registry.example.com: no such host
dial tcp 52.x.x.x:443: i/o timeout
# Registry uses a certificate the node does not trust
x509: certificate signed by unknown authorityCause 1: the tag or repository is wrong
This is the most common cause, and the fastest to rule out. Check that the exact tag exists in the repository the pod names. For ECR, list the images that the registry actually holds:
aws ecr describe-images \
--repository-name my-api \
--image-ids imageTag=v1.4.9 \
--region eu-west-1
# List recent tags to compare against what the deployment asks for
aws ecr describe-images --repository-name my-api \
--query 'sort_by(imageDetails,&imagePushedAt)[-5:].imageTags[]' \
--region eu-west-1Watch for a tag that was built but not pushed, a branch name that contains a slash, and a repository name that differs by one letter between the pipeline and the workload.
Cause 2: ECR authentication on EKS
Pods in the same AWS account as the ECR repository pull through the node role. The EKS managed node group role needs the read-only ECR policy, AmazonEC2ContainerRegistryReadOnly. The kubelet uses the ECR credential provider that ships with EKS, so the token is refreshed without a secret in the cluster.
Check that the node role has the policy, then check the image is in the same region as the cluster:
# Node role used by the managed node group
aws eks describe-nodegroup --cluster-name my-cluster \
--nodegroup-name ng-default --region eu-west-1 \
--query 'nodegroup.nodeRole'
aws iam list-attached-role-policies \
--role-name eks-node-role \
--query 'AttachedPolicies[].PolicyName'A pull from a different account, or from a registry in another region, needs a repository policy that allows the node role, or a secret. Avoid a static ECR password in a secret. It expires after 12 hours and the pods then fail again on the next pull.
Cause 3: a private registry needs an imagePullSecret
For a registry that is not ECR, such as GitHub Container Registry, Docker Hub with a private repository, or a self-hosted registry, create a docker-registry secret in the same namespace as the pod. Read the token from an environment variable so it stays out of your shell history.
kubectl create secret docker-registry regcred \
--namespace prod \
--docker-server=ghcr.io \
--docker-username="$REGISTRY_USER" \
--docker-password="$REGISTRY_TOKEN"
# Attach it to the service account the pod runs as
kubectl patch serviceaccount default -n prod \
-p '{"imagePullSecrets":[{"name":"regcred"}]}'
# Recreate the pods so they pick up the secret
kubectl rollout restart deployment/my-api -n prodA secret in another namespace is not visible to the pod. The imagePullSecrets name must match a secret in the pod namespace.
Cause 4: the nodes cannot reach the registry
Nodes in a private subnet with no NAT route and no VPC endpoints cannot reach ECR, so the pull times out. Test from a pod in the same subnet as the nodes. Use a throwaway pod so you test the real network path:
kubectl run pull-test --rm -it --restart=Never \
--image=busybox:1.36 -n prod -- \
nslookup 111122223333.dkr.ecr.eu-west-1.amazonaws.com
# If DNS works but the pull times out, check the route and the endpoints
aws ec2 describe-vpc-endpoints --region eu-west-1 \
--query 'VpcEndpoints[].ServiceName'For ECR from private subnets, add the ECR API, ECR Docker and S3 endpoints, or a NAT gateway. A private registry outside AWS needs the node egress rules that allow port 443.
Retry the pull after the fix
The kubelet retries on its own, but the back-off can be several minutes. Recreate the pod to retry at once:
kubectl delete pod my-api-7c9f -n prod
kubectl get pods -n prod -wIf the Deployment still points at the old image, the new pod fails the same way. Fix the image reference in the Deployment, then roll it out again.
Prevention checklist
- Tag images with an immutable value such as the git SHA, and keep the tag in the Deployment spec.
- Confirm the tag exists in the registry as part of the pipeline, before the deploy step.
- Use the node role for same-account ECR. Use a secret only for registries outside it, and rotate it.
- Keep an ECR VPC endpoint set in every private subnet that runs nodes.
What to check after the fix
The pod should reach Running with a ready container. kubectl get pods shows READY 1/1 and no restarts from the pull. If the pod shows ImagePullBackOff again, read the new event text, because the cause has changed.
How SeaGit handles this
SeaGit runs on AWS only, and the clusters it creates live in your AWS account. When you set registry credentials for an app, SeaGit stores them as a Kubernetes image pull secret named after the deployment and attaches it to the pod spec. You do not create the secret by hand.
When a deploy fails because a pod cannot pull its image, SeaGit's pod diagnosis reports ImagePullBackOff or ErrImagePull and names the image and the pod. The message points to the tag and the registry access. It does not replace the events above, so keep them handy when the tag is correct.
The network and IAM steps in this guide still apply inside your account. SeaGit does not change your node role, your VPC endpoints or your registry policy.