Troubleshooting
Route 53 Orphaned Records: Find and Delete Safely
You deleted the app. The Ingress is gone, the load balancer is gone, and the Route 53 zone still lists the name. Worse, the name keeps pointing at an address that no longer serves anything, and a new app with the same name inherits the stale record. This guide explains why ExternalDNS leaves those records, how to list them, and how to delete them without tripping Route 53’s matching rules.
By Muhammad Soliman, Founder at SeaGit
·
Key takeaways
- ExternalDNS with --policy=upsert-only creates and updates records but never deletes them. Every teardown under that policy leaves records behind.
- Route 53 matches a DELETE on every value. An alias record needs its AliasTarget repeated, or the change is rejected.
- Find orphans by comparing alias targets with live load balancers, and owner TXT records with live Ingress and Service objects.
- Delete from an exact copy of the stored record, in one change batch, after a snapshot and a review of the batch.
- Prevent the buildup with a unique owner ID per cluster, the right policy for each zone, and a teardown order that removes records before their controller.
Why records outlive their workloads
An orphaned record is a DNS name that still exists after the thing it pointed at is gone. ExternalDNS produces them for four reasons. The policy never deletes. The owner ID changed, so the new controller does not recognise the old records. The controller was removed before the records were cleaned up. Or the load balancer went first, which leaves an alias that points at nothing. None of these is a Route 53 bug. Each one is a gap in the teardown.
Teardown order decides the outcome
Under upsert-only, no order of deletion cleans up the records, because the controller does not delete them at any step. Under sync, the order still matters. If the controller is removed before the Ingress is deleted, nothing is left to notice the source is gone. If the load balancer is deleted first, the alias is left pointing at a target that no longer exists, and nothing in the zone records that the source is gone.
The safe order is to remove the source while the controller still runs, let the controller delete its own records, and only then remove the controller and the infrastructure behind it.
Why alias deletes fail
The Route 53 reference states the rule plainly: to delete a resource record set, you must specify all the same values that you specified when you created it. An alias record is created with an AliasTarget block, so a DELETE that only names the record and its type does not match it. The API then reports that it expected one of the value forms and found none.
# Illustrative output, not captured from a live zone
An error occurred (InvalidChangeBatch) when calling the ChangeResourceRecordSets operation:
Expected exactly one of [AliasTarget, all of [TTL, and ResourceRecords], or TrafficPolicyInstanceId], but found noneThe fix is to copy the stored object. A correct DELETE repeats the alias target, and includes SetIdentifier, Region, Weight and the other routing fields when the record has them. A simple record needs its TTL and ResourceRecords.
{
"Action": "DELETE",
"ResourceRecordSet": {
"Name": "old-app.example.com.",
"Type": "A",
"AliasTarget": {
"HostedZoneId": "Z0ELBZONEEXAMPLE",
"DNSName": "dualstack.k8s-old-app-0123456789.us-east-1.elb.amazonaws.com.",
"EvaluateTargetHealth": false
}
}
}{
"Action": "DELETE",
"ResourceRecordSet": {
"Name": "old-app.example.com.",
"Type": "A"
}
}Find the orphans
Start with a snapshot. Nothing else in this guide is safe without one, because a change batch cannot be undone by Route 53. Then run two independent checks and compare the results. The first checks alias targets against live load balancers. The second checks owner TXT records against live Ingress and Service objects.
ZONE=Z0123456789EXAMPLE
# Snapshot every record in the zone before changing anything
aws route53 list-resource-record-sets \
--hosted-zone-id "$ZONE" \
--output json > zone-before.json
# Count records and owner TXT entries for a baseline
jq '[.ResourceRecordSets[] | .Type] | group_by(.) | map({type: .[0], count: length})' zone-before.json# Alias targets in the zone, normalised to bare load balancer names
jq -r '.ResourceRecordSets[] | select(.AliasTarget) | .AliasTarget.DNSName' zone-before.json \
| sed -e 's/^dualstack\.//' -e 's/\.$//' | tr 'A-Z' 'a-z' | sort -u > alias-targets.txt
# Live Application and Network Load Balancers in the account and region
aws elbv2 describe-load-balancers --query 'LoadBalancers[].DNSName' --output text \
| tr '\t' '\n' | tr 'A-Z' 'a-z' | sort -u > live-lbs.txt
# Alias targets with no live load balancer behind them: candidates only
comm -23 alias-targets.txt live-lbs.txtThe second check finds the records the first one misses, such as a name that points at a live load balancer shared with another app. ExternalDNS writes an owner TXT record that names the source object, so you can check that object directly.
# Owner TXT records name the source object: resource=<kind>/<namespace>/<name>
jq -r '.ResourceRecordSets[] | select(.Type == "TXT") | .ResourceRecords[].Value' zone-before.json \
| grep -o 'external-dns/resource=[^,"]*' | sort -u > owners.txt
# Keep the ones whose Ingress or Service no longer exists
while read -r res; do
r=${res#external-dns/resource=}
kind=${r%%/*}; rest=${r#*/}; ns=${rest%%/*}; name=${rest#*/}
case "$kind" in
ingress|service)
kubectl -n "$ns" get "$kind" "$name" >/dev/null 2>&1 || echo "orphan owner: $r" ;;
esac
done < owners.txtThe orphan set is the intersection you trust, not the union. A record with a missing owner and a dead alias target is safe to review first. A record with only one signal needs a closer look before it goes.
Clean up with exact-match deletes
Build the change batch from the snapshot, not from memory. Copy each matched record object whole into a DELETE change. That keeps the AliasTarget, SetIdentifier and routing fields exactly as stored. Put the address record and its owner TXT record in the same batch. Exclude the zone apex NS and SOA records, which the zone needs.
# Pick the exact records by name, keep the owner TXT with its address record
jq --arg n1 "old-app.example.com." --arg n2 "_owner.old-app.example.com." '{
Comment: "remove orphaned records for old-app",
Changes: [
.ResourceRecordSets[]
| select(.Name == $n1 or .Name == $n2)
| select(.Type != "NS" and .Type != "SOA")
| {Action: "DELETE", ResourceRecordSet: .}
]
}' zone-before.json > delete.json
# Review the batch before sending it
jq '.Changes[] | {Action, Name: .ResourceRecordSet.Name, Type: .ResourceRecordSet.Type}' delete.json# Apply the batch. Route 53 accepts it all or nothing.
aws route53 change-resource-record-sets \
--hosted-zone-id "$ZONE" \
--change-batch file://delete.json
# Wait for the change to reach INSYNC
aws route53 get-change --id /change/C0123456789EXAMPLEWork in small batches, one application at a time, and re-run the comparison after each. If a delete reports that a record was not found, the record was already removed, so re-list the zone before retrying rather than sending a blind second delete.
Do not hand-write the batch from an address and a name. A record written from memory is the most common cause of the InvalidChangeBatch error above.
Prevention
- Pick the policy per zone. Use
--policy=syncin a zone that one controller owns, so deletes happen automatically. Use upsert-only in a shared zone and budget for a cleanup job. - Give each cluster a stable owner ID. The TXT registry documentation says clusters that share a zone need different owner IDs, and that an owner ID must not change for the life of the deployment. Changing it makes the old records invisible to the new controller.
- Tear down in the right order. Remove the Ingress or Service first, wait for the controller to act, then remove the controller and the load balancer.
- Audit on a schedule. Run the two checks from this guide, in a dry run, and review the output. An audit that is never read does nothing.
Orphans build up quietly. A zone with a few hundred dead names still answers queries, and a future app that reuses a name inherits the stale record. The cheapest time to clean up is at teardown, while the owner TXT record still says which controller wrote it.
How SeaGit handles this
The cluster worker on its release branch does three things about this problem. Its Route 53 delete paths list the record first and delete by its stored shape, which covers alias records. When two teardowns race for the same shared record, a delete that fails only because the record is already gone is treated as success, so the teardown does not report a failure it does not have. And a separate DNS reaper deletes only the records that the TXT registry proves its controller owns. The external-dns controllers still run with upsert-only, so the reaper is the part that removes the records.
If you run your own controller, the same rules apply. The DNS documentation describes how zones and domains attach to clusters, and the Route 53 application guide covers the application side.