Kubernetes troubleshooting
Unbound Immediate PersistentVolumeClaims: The Fix
A pod stays Pending with one of two scheduler messages. Both come from the same place: a PersistentVolumeClaim that is bound to a volume the pod cannot reach. This guide shows how to read the events, find the zone of the volume, and fix the binding mode and node placement so it does not happen again.
By Muhammad Soliman, Founder at SeaGit
·
Key takeaways
- pod has unbound immediate PersistentVolumeClaims means the claim has no volume yet. Check the default StorageClass and the CSI driver first.
- volume node affinity conflict means the volume exists in one zone, and no node in that zone has room for the pod.
- On AWS, EBS volumes are zonal. Use volumeBindingMode WaitForFirstConsumer so the volume is created where the pod lands.
- Changing the binding mode does not move an existing volume. Snapshot and restore it into the zone you need.
- For stateful workloads, keep enough capacity in every zone that holds a volume. Scaling a node group can move capacity out of that zone.
What the two errors mean
The kube-scheduler reports pods that it cannot place. It checks the storage a pod needs as part of that decision. Two messages come from that check, and they describe different states of the same claim.
pod has unbound immediate PersistentVolumeClaims means the claim uses the Immediate binding mode, which is the Kubernetes default when the field is unset. Immediate binding asks the provisioner to create the volume as soon as the claim exists. If the claim is still unbound at scheduling time, the scheduler will not place the pod. The volume may never be created because no default StorageClass exists, the CSI driver is not running, or the IAM role of the driver cannot create volumes.
volume node affinity conflict means the claim is bound. The PersistentVolume already exists, and its nodeAffinity lists the nodes that can use it. The scheduler then finds that none of those nodes can run the pod, because of capacity, taints, or a zone mismatch. On AWS the zone mismatch is the common case, and it is the one that appears after a node group rebalances across Availability Zones.
Illustrative scheduler output for each state is shown below. The node counts and the exact wording depend on your Kubernetes version, so use the text in your own events as the source of truth.
# Illustrative output, not copied from a cluster
Warning FailedScheduling 0/6 nodes are available: pod has unbound immediate PersistentVolumeClaims.
# Illustrative output for a bound zonal volume
Warning FailedScheduling 0/6 nodes are available: 3 node(s) had volume node affinity conflict.Read the events
Start with the pod. Its Events section names the failing check, and the claim events show whether provisioning ever started. Run these first, with the namespace and pod name from your workload:
kubectl describe pod my-db-0 -n datakubectl get events -n data --sort-by=.lastTimestamp
kubectl describe pvc data-my-db-0 -n dataIf the claim event says the volume is waiting for a first consumer, the claim is fine and the pod has not been scheduled yet. If the event says the provisioner failed, the storage layer is the problem and the rest of this guide applies. If there are no claim events at all, check that the StorageClass named by the claim exists.
Diagnose the claim and the volume
Once the claim is bound, the PersistentVolume tells you where the data lives. Print its node affinity and compare the zone with the zones of your nodes:
# Find the PV bound to the claim, then print its node affinity
PV=$(kubectl get pvc data-my-db-0 -n data -o jsonpath='{.spec.volumeName}')
kubectl get pv "$PV" -o jsonpath='{.spec.nodeAffinity.required.nodeSelectorTerms}{"\n"}'
kubectl get pv "$PV" -o yaml | grep -A4 nodeAffinityNext, check the StorageClass. The first question is whether one class is marked as the default. The second is which binding mode each class uses.
# Is there a default StorageClass, and what binding mode does each one use?
kubectl get storageclass
kubectl get storageclass -o yaml | grep -E 'name:|is-default-class|volumeBindingMode'# Which zone is each node in? Compare with the zone in the PV affinity.
kubectl get nodes -L topology.kubernetes.io/zone
# Is the EBS CSI controller running?
kubectl get pods -n kube-system | grep ebs-csiThree answers cover most cases. No default class and no class named in the claim means nothing can provision the volume. A default class with Immediate binding means the volume was created before the scheduler chose a node. A PV zone with no spare nodes means the pod is pinned to a zone without capacity.
The four common causes
- No default StorageClass. A claim without storageClassName uses the default class. If no class carries the default annotation, the claim waits and the pod reports unbound immediate claims. Some clusters have a class but lost the annotation when it was recreated.
- Immediate binding on zonal storage. The volume is created in the zone chosen at claim time, before any pod is placed. Any later pod that cannot run in that zone gets the affinity conflict. This is the cause behind most reports of the message on EBS.
- Zonal volumes on a multi-zone node group. An Auto Scaling group that spans zones can move instances between zones during a rebalance. The pods on those instances are rescheduled, and a pod with a volume in the old zone has nowhere to run if that zone is now short of capacity. Node groups that scale to zero in one zone make this worse.
- The EBS CSI driver is missing or cannot create volumes. On Amazon EKS, the driver runs as an add-on and needs an IAM role with the AmazonEBSCSIDriverPolicy managed policy, or a custom policy that grants the same EC2 volume actions. Without it, provisioning fails and the claim stays unbound.
Only the first and the last cause are configuration errors you can see in one command. The second and third are about timing and placement, and they appear only after the first scheduling decision has been made. That is why the same workload can work for months and then fail after a node replacement.
Fix: WaitForFirstConsumer
The fix for zonal storage is to delay volume creation until the scheduler has picked a node. The StorageClass field that does this is volumeBindingMode: WaitForFirstConsumer. The Kubernetes storage documentation describes it for topology-constrained backends, which is the case for an EBS volume. With this mode the claim stays Pending until a pod uses it, and the volume is created in the zone of the chosen node.
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
name: ebs-gp3-wffc
provisioner: ebs.csi.aws.com
parameters:
type: gp3
csi.storage.k8s.io/fstype: ext4
volumeBindingMode: WaitForFirstConsumer
reclaimPolicy: Retain
allowVolumeExpansion: truePoint the workload at the new class. For a StatefulSet, set the class in volumeClaimTemplates. Each replica gets its own claim from that template, so each replica’s volume is created in the zone where its pod lands:
apiVersion: apps/v1
kind: StatefulSet
metadata:
name: my-db
namespace: data
spec:
serviceName: my-db
replicas: 3
selector:
matchLabels:
app: my-db
template:
metadata:
labels:
app: my-db
spec:
containers:
- name: db
image: postgres:16
volumeClaimTemplates:
- metadata:
name: data
spec:
accessModes: ["ReadWriteOnce"]
storageClassName: ebs-gp3-wffc
resources:
requests:
storage: 20GiThe Kubernetes documentation notes that once WaitForFirstConsumer is used, allowedTopologies is no longer needed in most situations. It remains available if you must keep a class in one zone:
# Optional: restrict a class to one zone. Usually not needed with WaitForFirstConsumer.
allowedTopologies:
- matchLabelExpressions:
- key: topology.kubernetes.io/zone
values:
- us-east-1aFix a volume that is already in the wrong zone
A new StorageClass does not change a volume that already exists. You have two ways to get the data into a zone where pods can run. The first is to make sure the zone of the volume has capacity, so the pod can run next to the data. The second is to move the data.
To move it, take a snapshot with the EBS CSI driver, then create a new claim from the snapshot. The new claim uses the WaitForFirstConsumer class, so it binds in the zone of the next pod. The snapshot path needs the CSI snapshot controller and snapshot CRDs installed. The Amazon EKS guide lists the snapshot components as a prerequisite for the snapshot feature of the EBS driver.
# Snapshot the old volume, then restore it into a new claim.
# The new claim binds in the zone the scheduler chooses.
apiVersion: snapshot.storage.k8s.io/v1
kind: VolumeSnapshot
metadata:
name: data-my-db-0-snap
namespace: data
spec:
volumeSnapshotClassName: ebs-snapclass
source:
persistentVolumeClaimName: data-my-db-0
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: data-my-db-0-restored
namespace: data
spec:
storageClassName: ebs-gp3-wffc
accessModes: ["ReadWriteOnce"]
resources:
requests:
storage: 20Gi
dataSource:
name: data-my-db-0-snap
kind: VolumeSnapshot
apiGroup: snapshot.storage.k8s.ioStop the pod before taking the final snapshot of a database. A snapshot of a running volume can capture a state that the application cannot start from. Check the restore with the application’s own start-up and integrity checks before you remove the old claim.
Fix node placement for zonal volumes
Binding mode decides where a new volume is created. Node placement decides whether the pod can use it later. Keep the two aligned:
- Keep spare capacity in every zone that holds a volume. If a zone has no room for the pod, the pod stays Pending even though the volume is healthy. Node groups with a maximum per zone must leave room for the zone of each stateful pod.
- Spread replicas, but know the limit. A topology spread constraint keeps replicas in different zones. It cannot move an existing replica, so a replica pinned to a zone can still wait. Use it with WaitForFirstConsumer, not as a substitute for it.
- Do not rebalance stateful pools blindly. Before scaling a node group, compare the zone of each node with the zones of your PVs. Suspending rebalancing is an option for pools that run zonal workloads, but it does not replace capacity planning.
# In the pod template of the StatefulSet
topologySpreadConstraints:
- maxSkew: 1
topologyKey: topology.kubernetes.io/zone
whenUnsatisfiable: ScheduleAnyway
labelSelector:
matchLabels:
app: my-dbA worked diagnosis
The walk-through below is illustrative. The names, zones and counts are invented to show the order of the checks, and they are not output from a real cluster. It follows a common pattern: a database replica that restarts after a node group change.
Step one is the pod. Its Events show the scheduler message volume node affinity conflict, so the claim is bound and the problem is placement. Step two reads the PV. Its node affinity names zone us-east-1a. Step three lists the nodes with their zones. Step four checks the Auto Scaling group for that zone. In the illustrative case the group for zone a is at its maximum size, while the groups for zones b and c have room. The pod can only run in zone a, so it waits.
There are two ways out. Raising the maximum of the zone-a group gives the autoscaler room to add a node there. Alternatively, the data can be moved to zone b or c with the snapshot method above. The first is faster and keeps the data in place. The second is the right choice when zone a is being retired from the pool. Either way, record which zone each replica uses, so the next change to the node group is checked against it.
Per-zone node pools for StatefulSets
A single node group that spans three zones is convenient, but it hides the zone of each replica. A layout that suits zonal storage is one node group per zone, each with its own minimum and maximum. Every zone then has explicit capacity, and a StatefulSet replica pinned to a zone can always find room there.
The trade-off is more node groups to manage, and more chances for one zone to run short. Set each group’s maximum to cover the stateful pods in its zone plus normal headroom. Also check the autoscaler settings. Options that balance similar node groups across zones can move capacity away from the zone a volume needs, so test them with a stateful workload before you rely on them.
Stateless workloads do not need this layout. They can use the spread constraints shown above and let the scheduler choose any zone with capacity. Keep the zonal layout for the pods that own a volume.
Prevention checklist
- Exactly one StorageClass carries the default annotation, and it is the class you intend to use.
- Zonal block storage uses WaitForFirstConsumer. Immediate is reserved for storage that is reachable from every node.
- The EBS CSI add-on is installed, healthy, and has the IAM policy it needs.
- Stateful pools keep capacity in every zone that already holds a volume.
- A test pod is rescheduled to another zone before a node group change reaches production.
- Alerts fire on pods Pending longer than a few minutes, with the scheduler message in the alert.
How SeaGit handles this
SeaGit provisions EKS clusters in your AWS account. The EBS add-on creates the storage classes, and the defaults follow the rules above. The cluster-worker EBS add-on template on the release branch defines four StorageClasses: ebs-csi-im-deland ebs-csi-im-rtn use Immediate binding, and ebs-csi-slow-del and ebs-csi-slow-rtn use WaitForFirstConsumer. The Delete and Retain variants differ only in reclaim policy.
Stateful workloads should use one of the WaitForFirstConsumer classes. The limits in this article still apply. A volume that was created in one zone stays there, and a zone with no spare nodes still blocks the pod. Check the zone of the volume and the node capacity in that zone before you change a node group. For the cluster model, see the clusters documentation.