Troubleshooting

Route 53 Orphaned Records: Find and Delete Safely

You deleted the app. The Ingress is gone, the load balancer is gone, and the Route 53 zone still lists the name. Worse, the name keeps pointing at an address that no longer serves anything, and a new app with the same name inherits the stale record. This guide explains why ExternalDNS leaves those records, how to list them, and how to delete them without tripping Route 53’s matching rules.

By , Founder at SeaGit

·

Key takeaways

  • ExternalDNS with --policy=upsert-only creates and updates records but never deletes them. Every teardown under that policy leaves records behind.
  • Route 53 matches a DELETE on every value. An alias record needs its AliasTarget repeated, or the change is rejected.
  • Find orphans by comparing alias targets with live load balancers, and owner TXT records with live Ingress and Service objects.
  • Delete from an exact copy of the stored record, in one change batch, after a snapshot and a review of the batch.
  • Prevent the buildup with a unique owner ID per cluster, the right policy for each zone, and a teardown order that removes records before their controller.

Why records outlive their workloads

An orphaned record is a DNS name that still exists after the thing it pointed at is gone. ExternalDNS produces them for four reasons. The policy never deletes. The owner ID changed, so the new controller does not recognise the old records. The controller was removed before the records were cleaned up. Or the load balancer went first, which leaves an alias that points at nothing. None of these is a Route 53 bug. Each one is a gap in the teardown.

Teardown order decides the outcome

Under upsert-only, no order of deletion cleans up the records, because the controller does not delete them at any step. Under sync, the order still matters. If the controller is removed before the Ingress is deleted, nothing is left to notice the source is gone. If the load balancer is deleted first, the alias is left pointing at a target that no longer exists, and nothing in the zone records that the source is gone.

The safe order is to remove the source while the controller still runs, let the controller delete its own records, and only then remove the controller and the infrastructure behind it.

Teardown order decides whether Route 53 records surviveTop row, leaking order: the controller is removed first, then the load balancer, then the Ingress. Records remain because nothing is left to delete them. Bottom row, safe order: the Ingress is removed while the controller still runs, the controller deletes its records under sync policy, and only then is the controller removed.Leaks: controller removed before the records goRemove controllerno reconciler leftDelete load balanceralias target goneDelete Ingressnothing deletes recordsSafe: delete the source, let the controller clean up, then remove itDelete Ingresscontroller still runningController deletessync policy, owned onlyRemove controllerthen the load balancerUnder upsert-only the middle box never happens, so the orphan step is the same whatever the order.
Top row: the controller goes first and the records stay. Bottom row: the source goes while the controller is still running, so a sync policy can delete the records it owns.

Why alias deletes fail

The Route 53 reference states the rule plainly: to delete a resource record set, you must specify all the same values that you specified when you created it. An alias record is created with an AliasTarget block, so a DELETE that only names the record and its type does not match it. The API then reports that it expected one of the value forms and found none.

text
# Illustrative output, not captured from a live zone
An error occurred (InvalidChangeBatch) when calling the ChangeResourceRecordSets operation:
Expected exactly one of [AliasTarget, all of [TTL, and ResourceRecords], or TrafficPolicyInstanceId], but found none
Illustrative error text in the shape Route 53 returns for a name-and-type-only alias delete.

The fix is to copy the stored object. A correct DELETE repeats the alias target, and includes SetIdentifier, Region, Weight and the other routing fields when the record has them. A simple record needs its TTL and ResourceRecords.

json
{
  "Action": "DELETE",
  "ResourceRecordSet": {
    "Name": "old-app.example.com.",
    "Type": "A",
    "AliasTarget": {
      "HostedZoneId": "Z0ELBZONEEXAMPLE",
      "DNSName": "dualstack.k8s-old-app-0123456789.us-east-1.elb.amazonaws.com.",
      "EvaluateTargetHealth": false
    }
  }
}
An alias DELETE that matches the stored record. The HostedZoneId is the load balancer zone ID, as stored, and the values are placeholders.
json
{
  "Action": "DELETE",
  "ResourceRecordSet": {
    "Name": "old-app.example.com.",
    "Type": "A"
  }
}
A DELETE that will be rejected. It names the record but omits the value that identifies it.
A DELETE must repeat the stored record's values, including the alias targetLeft column: the stored alias record with name, type, alias target and evaluate-target-health. Right column: a delete request that carries only name and type. Route 53 rejects that request. The lower right shows the same delete with every stored field copied, which Route 53 accepts.Stored record (what exists)Name: app.example.com. AAliasTarget.DNSName: dualstack.k8s-...elb...AliasTarget.HostedZoneId: Z35SXDOTRQ7X7KAliasTarget.EvaluateTargetHealth: false(zone ID shown is illustrative)Name and type only: rejectedAction: DELETEName: app.example.com.Type: AInvalidChangeBatch: Expected exactly one of[AliasTarget, ...], but found noneCopy the stored record: acceptedAction: DELETE + the full ResourceRecordSetInclude SetIdentifier and routing fields if the record has them.Route 53 matches a delete on every value you specify.So list the record first and send that exact object back.
The stored record and the delete request must agree on every value. A request with only the name and type fails, and a copy of the stored object succeeds.

Find the orphans

Start with a snapshot. Nothing else in this guide is safe without one, because a change batch cannot be undone by Route 53. Then run two independent checks and compare the results. The first checks alias targets against live load balancers. The second checks owner TXT records against live Ingress and Service objects.

shell
ZONE=Z0123456789EXAMPLE

# Snapshot every record in the zone before changing anything
aws route53 list-resource-record-sets \
  --hosted-zone-id "$ZONE" \
  --output json > zone-before.json

# Count records and owner TXT entries for a baseline
jq '[.ResourceRecordSets[] | .Type] | group_by(.) | map({type: .[0], count: length})' zone-before.json
Snapshot the zone to a file. Keep it with the change record, because it is the only exact copy of the values.
shell
# Alias targets in the zone, normalised to bare load balancer names
jq -r '.ResourceRecordSets[] | select(.AliasTarget) | .AliasTarget.DNSName' zone-before.json \
  | sed -e 's/^dualstack\.//' -e 's/\.$//' | tr 'A-Z' 'a-z' | sort -u > alias-targets.txt

# Live Application and Network Load Balancers in the account and region
aws elbv2 describe-load-balancers --query 'LoadBalancers[].DNSName' --output text \
  | tr '\t' '\n' | tr 'A-Z' 'a-z' | sort -u > live-lbs.txt

# Alias targets with no live load balancer behind them: candidates only
comm -23 alias-targets.txt live-lbs.txt
Candidates are alias targets with no matching live load balancer. Check each one before you delete anything: the target may belong to another account, a different region, or a service that is not a load balancer.

The second check finds the records the first one misses, such as a name that points at a live load balancer shared with another app. ExternalDNS writes an owner TXT record that names the source object, so you can check that object directly.

shell
# Owner TXT records name the source object: resource=<kind>/<namespace>/<name>
jq -r '.ResourceRecordSets[] | select(.Type == "TXT") | .ResourceRecords[].Value' zone-before.json \
  | grep -o 'external-dns/resource=[^,"]*' | sort -u > owners.txt

# Keep the ones whose Ingress or Service no longer exists
while read -r res; do
  r=${res#external-dns/resource=}
  kind=${r%%/*}; rest=${r#*/}; ns=${rest%%/*}; name=${rest#*/}
  case "$kind" in
    ingress|service)
      kubectl -n "$ns" get "$kind" "$name" >/dev/null 2>&1 || echo "orphan owner: $r" ;;
  esac
done < owners.txt
Each owner whose Ingress or Service is missing is an orphan candidate. The check depends on kubectl access to every namespace the zone names.

The orphan set is the intersection you trust, not the union. A record with a missing owner and a dead alias target is safe to review first. A record with only one signal needs a closer look before it goes.

Clean up with exact-match deletes

Build the change batch from the snapshot, not from memory. Copy each matched record object whole into a DELETE change. That keeps the AliasTarget, SetIdentifier and routing fields exactly as stored. Put the address record and its owner TXT record in the same batch. Exclude the zone apex NS and SOA records, which the zone needs.

shell
# Pick the exact records by name, keep the owner TXT with its address record
jq --arg n1 "old-app.example.com." --arg n2 "_owner.old-app.example.com." '{
  Comment: "remove orphaned records for old-app",
  Changes: [
    .ResourceRecordSets[]
    | select(.Name == $n1 or .Name == $n2)
    | select(.Type != "NS" and .Type != "SOA")
    | {Action: "DELETE", ResourceRecordSet: .}
  ]
}' zone-before.json > delete.json

# Review the batch before sending it
jq '.Changes[] | {Action, Name: .ResourceRecordSet.Name, Type: .ResourceRecordSet.Type}' delete.json
The batch is built by name and copied from the snapshot. Read the review output before you send it.
shell
# Apply the batch. Route 53 accepts it all or nothing.
aws route53 change-resource-record-sets \
  --hosted-zone-id "$ZONE" \
  --change-batch file://delete.json

# Wait for the change to reach INSYNC
aws route53 get-change --id /change/C0123456789EXAMPLE
Apply and wait. A change batch is applied all or nothing, so a rejected batch leaves the zone as it was.

Work in small batches, one application at a time, and re-run the comparison after each. If a delete reports that a record was not found, the record was already removed, so re-list the zone before retrying rather than sending a blind second delete.

Do not hand-write the batch from an address and a name. A record written from memory is the most common cause of the InvalidChangeBatch error above.

Prevention

  • Pick the policy per zone. Use --policy=sync in a zone that one controller owns, so deletes happen automatically. Use upsert-only in a shared zone and budget for a cleanup job.
  • Give each cluster a stable owner ID. The TXT registry documentation says clusters that share a zone need different owner IDs, and that an owner ID must not change for the life of the deployment. Changing it makes the old records invisible to the new controller.
  • Tear down in the right order. Remove the Ingress or Service first, wait for the controller to act, then remove the controller and the load balancer.
  • Audit on a schedule. Run the two checks from this guide, in a dry run, and review the output. An audit that is never read does nothing.

Orphans build up quietly. A zone with a few hundred dead names still answers queries, and a future app that reuses a name inherits the stale record. The cheapest time to clean up is at teardown, while the owner TXT record still says which controller wrote it.

How SeaGit handles this

The cluster worker on its release branch does three things about this problem. Its Route 53 delete paths list the record first and delete by its stored shape, which covers alias records. When two teardowns race for the same shared record, a delete that fails only because the record is already gone is treated as success, so the teardown does not report a failure it does not have. And a separate DNS reaper deletes only the records that the TXT registry proves its controller owns. The external-dns controllers still run with upsert-only, so the reaper is the part that removes the records.

If you run your own controller, the same rules apply. The DNS documentation describes how zones and domains attach to clusters, and the Route 53 application guide covers the application side.

Frequently asked questions

Route 53 orphaned records: FAQ

Why does ExternalDNS not delete Route 53 records when I delete a Service or Ingress?

Because of the policy. With --policy=upsert-only the controller creates and updates records but never deletes them. The ExternalDNS AWS tutorial says to remove leftover records by hand under that policy. Switch to --policy=sync only in a zone that one controller owns.

How do I delete a Route 53 alias record?

Read the existing record first, then send a DELETE change that repeats every value from that record, including the AliasTarget with DNSName, HostedZoneId and EvaluateTargetHealth. A DELETE that carries only the name and type is rejected with an InvalidChangeBatch error.

What is an orphaned Route 53 record?

A record that still exists after the workload it served is gone. Typical signs are an alias pointing at a load balancer that no longer exists, and a TXT owner record whose resource field names an Ingress or Service that is not in the cluster any more.

Is it safe to delete a record that still resolves?

Confirm first that nothing serves the name. A record that resolves is not automatically in use, because the name may still point at a retired address. Check the target, the owner TXT record and the application owner, then delete the whole group together in one change batch.

Can a change batch remove several records at once?

Yes. Route 53 treats a change batch as transactional. Either every change succeeds, or none do. Put the address record and its owner TXT record in the same batch so the zone never holds one without the other.

How do I stop orphans from building up again?

Give each controller a stable, unique owner ID per cluster, avoid changing owner IDs on a live deployment, tear down in the right order, and run a regular audit that compares owner TXT records with live Ingress and Service objects.

Sources

Checked 10 October 2026.

  1. ChangeResourceRecordSets (Amazon Route 53 API Reference) — DELETE must repeat the values of the record, and a change batch is applied all or nothing
  2. ResourceRecordSet (Amazon Route 53 API Reference) — the fields that identify a record, including SetIdentifier and the alias target
  3. AWS tutorial (ExternalDNS documentation) — the upsert-only behaviour and the instruction to remove leftover DNS records by hand
  4. The TXT registry (ExternalDNS documentation) — owner IDs per cluster in a shared zone, and the rule that prefixes cannot change after deploy

Read next