Troubleshooting
CoreDNS CrashLoopBackOff on EKS
CoreDNS is the cluster’s DNS service. When its pods crash or stay Pending, every service lookup fails, and the error often shows up in application logs first. This guide sorts the causes by what kubectl shows, then gives the AWS CLI steps for the EKS add-on.
By Muhammad Soliman, Founder at SeaGit
·
Key takeaways
- Read the STATUS first. Pending means the pod was never scheduled, so it has no logs. CrashLoopBackOff means the container started and exited.
- A message that starts with Loop and ends with detected means a forwarding loop. The loop plugin stops CoreDNS on purpose, so fix the Corefile or the resolver file it forwards to.
- On EKS, CoreDNS is usually a managed add-on. Check its status and health with aws eks describe-addon before you change anything inside the cluster.
- Match the add-on version to your Kubernetes version using aws eks describe-addon-versions, then update with an explicit --resolve-conflicts policy.
What CoreDNS does in an EKS cluster
Pods reach Services by name, such as my-service.my-namespace.svc.cluster.local. The resolver in each pod sends those names to the cluster DNS service, which runs CoreDNS in the kube-system namespace. The CoreDNS pods read their behaviour from a ConfigMap named coredns, the Corefile. Names outside the cluster domain are forwarded to upstream resolvers listed in the Corefile.
When CoreDNS is unhealthy, most failures look like application errors: timeouts, unknown host errors, or retries that never finish. Start from the cluster side. The application logs will not tell you that the DNS pods are the cause.
Managed or self-managed
On Amazon EKS, CoreDNS can be an Amazon EKS add-on, which AWS lists by the name coredns, or a self-managed installation that you update yourself. The managed page says the version table can differ from the self-managed versions. Find out which one you run before you change anything:
# What EKS thinks is installed: status, version and health issues
aws eks describe-addon \
--cluster-name my-cluster \
--addon-name coredns \
--query 'addon.{status:status,version:addonVersion,issues:health.issues}'
# Which CoreDNS versions can run on this cluster's Kubernetes version
aws eks describe-addon-versions \
--addon-name coredns \
--kubernetes-version 1.33 \
--query 'addons[].addonVersions[].addonVersion'If describe-addon returns an addon block, the add-on is managed by EKS. If it returns a not-found error, the CoreDNS Deployment is self-managed, and you update it with the manifests or Helm chart that installed it.
The two types need different fixes. A managed add-on takes its version and its conflict policy from the EKS API, so the steps in the update section apply. A self-managed CoreDNS takes its image and its Corefile from whatever installed it. Changing the Deployment by hand on a self-managed install works until the next apply overwrites it, so put the change in the source that installed it.
Triage: Pending, CrashLoopBackOff or OOMKilled
List the pods with their node placement. The k8s-app=kube-dns label is the one the CoreDNS pods carry in a standard EKS install.
# Where are the CoreDNS pods, and are they Ready?
kubectl -n kube-system get pods -l k8s-app=kube-dns -o wide
# The Deployment: desired and available replicas
kubectl -n kube-system get deployment corednsThen read the STATUS column and the restart count. Three states matter. Figure 1 sorts them.
- Pending. The pod is waiting for a node. Read the events, not the logs.
- CrashLoopBackOff. The container runs and exits. Read the previous container’s logs, then the Corefile.
- OOMKilled. The kernel stopped the container for using more memory than its limit. The last state in
describesays so.
CrashLoopBackOff with a loop message
The CoreDNS loop plugin sends a random probe query to itself and counts how often it sees the probe come back. If it sees the probe more than twice, it assumes a forwarding loop and halts the server. The CoreDNS documentation describes this as a fatal error, because an endless loop would use memory and CPU until the host runs out.
The loop plugin’s troubleshooting section names two causes. The common one is CoreDNS forwarding to itself through a loopback address such as 127.0.0.1, ::1 or 127.0.0.53. The less common one is an upstream server that forwards back to CoreDNS. Logs look like this (illustrative):
# Illustrative log line. The exact text includes the zone and a random name.
[FATAL] plugin/loop: Loop (127.0.0.1:53 -> :53) detected for zone ".", see https://coredns.io/plugins/loop#troubleshootingFind the Corefile, then look at every forward line. A forward that points to a loopback address, or to a file that contains one, causes the loop.
# The Corefile lives in the coredns ConfigMap in kube-system.
kubectl -n kube-system get configmap coredns -o yaml
# Forwarding to a loopback address sends queries back to CoreDNS. This line causes the loop:
# forward . 127.0.0.1
# The default forwards to the node's resolver file, which must not point at CoreDNS:
# forward . /etc/resolv.confThe loop plugin tells you to check the file that forward reads. If it is /etc/resolv.conf, check the file on the node that the CoreDNS pod runs on. A nameserver line that points to 127.0.0.1 or to a cluster-internal address will loop. Fix the upstream, then restart the pods. A pod that has already crashed will not recover on its own.
Pending pods that never start
A Pending CoreDNS pod has not been scheduled. Kubernetes records the reason as an event on the pod, and kubectl describe prints it. The common reasons are a shortage of CPU or memory on the nodes, pod rules that no node satisfies, and no nodes at all. Scale the node group to zero and the pods have nowhere to run.
# Events explain Pending. For CrashLoopBackOff, read the Last State too.
kubectl -n kube-system describe pod <coredns-pod-name>
# Logs from the container that crashed, not the one restarting now
kubectl -n kube-system logs <coredns-pod-name> --previousAn event like the one below points to capacity, not to CoreDNS (illustrative):
# Illustrative output only. Your cluster will show different counts and reasons.
Events:
Type Reason Message
---- ------ -------
Warning FailedScheduling 0/2 nodes are available: 2 Insufficient memory.
preemption: 0/2 nodes are available.If the message names insufficient memory or CPU, add node capacity or free it up. If it names pod rules, such as an anti- affinity rule that the available nodes cannot satisfy, the rule needs more nodes in different zones, or a different rule. Two replicas with a hard rule that they run on different nodes need two nodes that can take them, not one. If the message names no nodes, the cluster has no capacity for the pods at all, and the fix is in the node group, not in CoreDNS.
OOMKilled and version mismatch
If the last state reads OOMKilled, the container hit its memory limit. Compare the limit on the Deployment with the memory use you see in the metrics for the pods. Raising the limit is the usual fix. If memory use keeps climbing after a raise, check the CoreDNS release notes for your version before you change more.
A version mismatch shows up after a cluster upgrade. The add-on keeps its old version, and the add-on health can report an issue. The managed CoreDNS page lists the latest version for each Kubernetes version. The CLI command in the previous section returns every version your cluster can use. Pick one from that list, not from another minor release.
Update the add-on safely
Updating a managed add-on replaces its configuration with the add-on’s own settings for any field that conflicts with a change you made by hand. The --resolve-conflicts flag sets what happens to those fields. The AWS add-on page shows OVERWRITE on the create command, and the update command takes the same policy.
# Update to a version from the list above. Choose the conflict policy after step 1.
aws eks update-addon \
--cluster-name my-cluster \
--addon-name coredns \
--addon-version v1.12.4-eksbuild.57 \
--resolve-conflicts PRESERVE
# Watch the add-on until it reports ACTIVE, then check the pods again.
aws eks describe-addon --cluster-name my-cluster --addon-name coredns --query 'addon.status'
kubectl -n kube-system rollout status deployment/corednsRun the update in a quiet window. DNS is on the path of almost every request, so a short period of slow or failed lookups during the update is expected, and it should be planned rather than discovered. Watch the CoreDNS pods during the rollout, and run the throwaway lookup from the previous section after the rollout finishes, not while it is still in progress.
Keep the Corefile change small when you make it. Save the current ConfigMap first, so a bad change has a reference to go back to. If the add-on reports a health issue after the update, read the issue text from describe-addon before you run another update. Two updates in a row, each with a different conflict policy, make it harder to tell which change caused a problem.
# Keep a copy of the current Corefile before any change
kubectl -n kube-system get configmap coredns -o yaml > coredns-corefile-backup.yaml
# After a Corefile change, restart the CoreDNS pods so they reload it
kubectl -n kube-system rollout restart deployment/corednsPrevent it
- Keep node capacity above the DNS requirement. DNS pods are small, but they need a node that has room. Check that the default node group can run the replicas after you scale it down, not only at its normal size.
- Keep the resolver chain free of loops. Do not forward CoreDNS to a loopback address, and do not point the node resolver at a CoreDNS Service address.
- Update the add-on with the cluster. Match the add-on version to the Kubernetes version after every cluster upgrade, and record the conflict policy you used.
- Keep the Corefile in version control. A saved copy of the Corefile, with the reason for each forward line, makes a loop easy to spot in review and quick to revert after an incident.
- Alert on CoreDNS restarts and Pending pods. A restart count that climbs, or a pod Pending for minutes, shows up as DNS errors in every application. Catch it in the DNS layer.
If DNS lookups fail intermittently rather than all at once, the cause is often a separate issue with the pod resolver, such as the ndots setting. The troubleshooting guide covers the general checks.
Check DNS from inside the cluster
Before you change CoreDNS, prove the failure from a pod. A one-off pod that resolves a cluster Service name tells you the resolver path works. A failure there, with the CoreDNS pods Ready, points at the Service, the network policy or the pod resolver. A failure with the CoreDNS pods not Ready points back to the triage above.
# Resolve a cluster Service name from a throwaway pod
kubectl run dns-check --rm -it --restart=Never --image=busybox:1.36 -- nslookup kubernetes.default.svc.cluster.local
# Resolve a public name the same way, to test the upstream forwarders
kubectl run dns-check --rm -it --restart=Never --image=busybox:1.36 -- nslookup example.comRun the same check twice: once for a name inside the cluster, and once for a public name such as an external API host. If only the public name fails, the upstream forwarders in the Corefile are the suspect, not the cluster domain. If both fail, look at the CoreDNS pods first. If only one namespace fails, check its NetworkPolicy objects before you change anything in kube-system.
Keep the output of each check. A timestamped result before and after a change is the only reliable way to tell whether the change fixed the problem or only changed its timing.
What a healthy add-on looks like
A healthy managed add-on reports status ACTIVE, has an addonVersion that matches your Kubernetes version in the AWS table, and returns an empty list for health issues. The CoreDNS Deployment shows its desired replicas as available, and the pods are Ready with a restart count that does not climb between checks. These are the four values to record before any change, so you can compare them afterwards.
A healthy Corefile has one forwarding path to upstream resolvers, a loop plugin, and no forward line that points to a loopback address. Compare your Corefile with the one you saved before any change. A forwarder that you added by hand is one way a healthy add-on starts looping after an edit.
How SeaGit handles this
SeaGit clusters on AWS install CoreDNS as an Amazon EKS managed add-on. The cluster module installs it with the VPC CNI and kube-proxy, so it follows the managed-add-on path in this guide. The limits:
- Version. The add-on version is chosen by SeaGit when the cluster is created or updated. Check it on your cluster with describe-addon, as shown above.
- Changes made outside SeaGit. A manual change to the CoreDNS add-on can conflict with the next update. Use the conflict steps in this guide, and check the cluster after each change.
- Limits of scope. This applies to the AWS clusters that SeaGit provisions in your account. Other clusters follow the same EKS steps, but SeaGit does not manage them.
For the cluster model, see the clusters documentation. For the application side of a failed lookup, see the applications documentation.