Kubernetes troubleshooting
Kubernetes Pod Stuck in Pending: Causes and Fixes
A Pending pod has not been placed on a node yet. The kube-scheduler writes the reason into the pod events, and that reason is the whole diagnosis. Read it, then either make the pod smaller or pickier, or give the scheduler a node that fits.
By Muhammad Soliman, Founder at SeaGit
·
Key takeaways
- Pending is a scheduling state. Nothing is wrong with the container yet, so do not look at its logs first.
- kubectl describe pod shows a line such as 0/3 nodes are available, followed by a count and a reason for each node group.
- Insufficient cpu or memory means the requests are larger than any free node. Lower the request, add capacity, or free space.
- Too many pods on EKS means the node reached its pod limit for its instance type. That limit depends on ENIs and IPs, so it is not fixed.
- Pending with unbound PersistentVolumeClaims is a storage problem, covered in its own guide.
Step 1: read the scheduler message
Check the pod events. The scheduler lists every node and the reason it was rejected, so one message usually covers the whole cluster.
kubectl get pod my-api-7b9d4c6f8-m2xqz -n prod -o wide
kubectl describe pod my-api-7b9d4c6f8-m2xqz -n prod | sed -n '/Events:/,$p'
# Recent scheduling events in the namespace
kubectl get events -n prod --field-selector reason=FailedScheduling \
--sort-by=.lastTimestamp | tail -10# Illustrative scheduler messages, not copied from a cluster
0/3 nodes are available: 3 Insufficient cpu.
0/3 nodes are available: 1 node(s) had untolerated taint {dedicated: batch}, 2 node(s) didn't match Pod's node affinity/selector.
0/2 nodes are available: 2 Too many pods.
0/3 nodes are available: 3 node(s) didn't have free ports for the requested pod ports.Cause 1: Insufficient cpu or memory
The scheduler adds up the requests of the pods already on a node, and compares the sum with the node allocatable capacity. A new pod fits only if its request fits in what is left. A request larger than any single node can never be scheduled, however empty the cluster is.
# Allocatable capacity per node
kubectl get nodes -o custom-columns=\
NAME:.metadata.name,\
CPU:.status.allocatable.cpu,\
MEMORY:.status.allocatable.memory,\
PODS:.status.allocatable.pods
# What is already requested on one node
kubectl describe node ip-10-0-12-34.eu-west-1.compute.internal | sed -n '/Allocated resources/,/Events/p'
# The request on the pod you are placing
kubectl get pod my-api-7b9d4c6f8-m2xqz -n prod \
-o jsonpath='{.spec.containers[*].resources.requests}{"\n"}'- Lower a request that is far above real use. Measure first, the same way as for memory limits in the OOMKilled guide.
- Add capacity in the node group that matches the pod, or use a larger instance type for that workload.
- Scale down or remove a workload that holds requests it does not use, such as an old preview deployment.
Cause 2: taints, selectors and affinity
A taint keeps pods off a node unless the pod has a matching toleration. A nodeSelector or node affinity keeps the pod on nodes with specific labels. A pod that asks for a label no node has stays Pending forever.
# Taints on each node
kubectl get nodes -o custom-columns=NAME:.metadata.name,TAINTS:.spec.taints
# Labels that the pod's nodeSelector or affinity expects
kubectl get pod my-api-7b9d4c6f8-m2xqz -n prod -o jsonpath='{.spec.nodeSelector}{"\n"}{.spec.affinity.nodeAffinity}{"\n"}'
kubectl get nodes --show-labelsFix the side that is wrong. If the node label is missing, add it to the node group. If the pod selector is a leftover from another environment, remove it. Add a toleration only when the taint is meant for this workload.
Cause 3: Too many pods on the node
Each node has a maximum pod count, set by its allocatable pods value. On EKS with the VPC CNI, that value follows the number of network interfaces and IP addresses the instance type supports. Small instances hit the limit with a few dozen pods, even when CPU and memory are free.
kubectl get nodes -L node.kubernetes.io/instance-type \
-o custom-columns=NAME:.metadata.name,PODS:.status.allocatable.pods
kubectl get pods -A --field-selector spec.nodeName=ip-10-0-12-34.eu-west-1.compute.internal --no-headers | wc -lTo see the limit for a given instance type, use the EKS max pods calculator. It shows the number from the ENI and IP rules and the effect of prefix delegation. Choose a larger instance type, or enable prefix delegation on the VPC CNI if your instance supports it, then roll the nodes so the new limit applies.
Cause 4: a hostPort is already in use
A pod that sets hostPort can run on only one node per port. A second pod that asks for the same port on the same node stays Pending with didn't have free ports. Remove hostPort and reach the pod through a Service, which handles the port for you. Only keep hostPort for a DaemonSet that must run one pod per node.
When the node group is at its maximum
Insufficient capacity can also mean the node group has reached its maximum size, so no new node can be added. The events do not always say so. Check the node group size limits against the current node count. Raising the maximum is a capacity decision, so check the cost and the limits of your account first.
What to check after the change
The pod should move from Pending to Running within a minute or two, and kubectl get events should stop showing FailedScheduling for it. If it stays Pending with a new message, the first fix has worked and a second cause is in place.
How SeaGit handles this
When a deploy waits on pods that cannot be scheduled, SeaGit's pod diagnosis reads the Unschedulable condition and the scheduler message, and reports the reason in the deploy status. You see Insufficient CPU or a similar reason, not a generic timeout.
Node group sizes, instance types and the resource requests of each app are set in SeaGit and applied to your cluster in your AWS account. Changing them is the fix for capacity causes. The scheduler rules in this guide apply in full, because SeaGit runs standard Kubernetes scheduling.