A senior-engineer guide to DNS delegation internals, automated edge-DNS synchronization with ExternalDNS, Let's Encrypt lifecycle automation via cert-manager, and the CLI toolkit to prove it all works end to end.
These two roles are the entire DNS data plane. Confusing them is the root cause of most "my DNS is wrong" tickets.
| Role | What it does | Examples | Key behavior |
|---|---|---|---|
| Recursive resolver | Walks the delegation chain on behalf of clients, starting at the root (.) and following referrals down to an authoritative server | 8.8.8.8, 1.1.1.1, your ISP resolver, CoreDNS in caching mode | Caches answers for the record TTL; answers are non-authoritative; a cache miss is not an error |
| Authoritative nameserver | Hosts the zone data for a domain and answers from that zone file only | Route 53's four nameservers for your hosted zone (e.g., ns-1024.awsdns-00.org) | Never caches; sets the AA (authoritative answer) flag; serves NXDOMAIN for names it owns that do not exist |
AA flag. When you need ground truth about a record, bypass every cache and query the authoritative nameservers directly (dig @ns-1024.awsdns-00.org example.com A — section 5).A hosted zone is a container for records about a domain (and its subdomains). Route 53 automatically provisions four nameservers (NS) and an SOA record at the zone apex the moment you create it.
example.com (Zone ID Z0123456789ABCDEFGHIJ), served by ns-1024.awsdns-00.org, ns-2048.awsdns-01.com, ns-3072.awsdns-02.co.uk, ns-4096.awsdns-03.net.example.com) must point at your four nameservers. Until the NS records match, your zone is a well-furnished room nobody can find.Subdomain delegation is pure NS-record mechanics. For a child zone dev.example.com (Zone ID Z0987654321KLMNOPQRST):
example.com) must contain identical NS records at the label dev.example.com. This is the delegation itself — there is no other signal.ns1.dev.example.com for dev.example.com) — otherwise resolvers deadlock trying to resolve the NS name through the zone being delegated. Route 53's nameservers live in awsdns-*.com/.net/.org/.co.uk, i.e. out-of-zone, so no glue is required.The parent zone's delegation records for dev.example.com:
dev.example.com. 172800 IN NS ns-2048.awsdns-01.com.
dev.example.com. 172800 IN NS ns-4096.awsdns-03.net.
dev.example.com. 172800 IN NS ns-1024.awsdns-00.org.
dev.example.com. 172800 IN NS ns-3072.awsdns-02.co.uk.
Full resolution path for app.example.com from a cold cache:
127.0.0.53) forwards to a recursive resolver.a.root-servers.net, 198.41.0.4) for example.com — gets a referral to the .com TLD (a.gtld-servers.net, 192.5.6.30).awsdns nameservers.ns-1024.awsdns-00.org — gets the authoritative A record: 203.0.113.42, TTL 300.| Type | Address family | Typical targets | Gotcha |
|---|---|---|---|
A | IPv4 (32-bit) | ALB/NLB IPs, EC2, EIP 203.0.113.42 | One A record set can hold multiple values (round-robin) |
AAAA | IPv6 (128-bit) | IPv6-only endpoints 2001:db8:85a3::8a2e:370:7334 | Clients with broken IPv6 connectivity fall back to A after a timeout — keep AAAA TTLs honest |
A CNAME is a pure alias: www.example.com CNAME k8s-prod-app-1234567890.eu-central-1.elb.amazonaws.com. RFC 1034 forbids a CNAME from coexisting with any other record at the same name. The zone apex (example.com) is permanently occupied by NS and SOA records — therefore a CNAME can never exist at the apex. It also can't point at another CNAME (chains are legal but slow and fragile), and an MX target may not be a CNAME.
An Alias is a Route 53 virtual record. It stores no answer of its own; at query time Route 53 resolves the target resource and returns its live A/AAAA values. Because it is virtual, it sidesteps every CNAME restriction: it works at the apex, it coexists with other records, it tracks IP changes of AWS resources automatically, it participates in health-check-weighted routing, and queries against AWS resources (ALB, CloudFront, S3 website endpoint, API Gateway, Global Accelerator, Elastic Beanstalk, or another record in the same zone) are free.
| Capability | CNAME | Route 53 Alias |
|---|---|---|
Usable at zone apex (example.com) | No — NS/SOA occupy the node | Yes |
| Targets | Any DNS name | AWS resources + records in the same zone only |
| Coexists with other records at same name | No | Yes (one Alias per record type) |
| Tracks target IP changes | No — relies on target TTL | Yes, at query time |
| Query cost | Standard rate | Free when targeting AWS resources |
| Visible in zone export/transfer | Yes | No — virtual record |
www.example.com gets a CNAME to it, but the apex example.com must be an Alias. This is exactly the pattern ExternalDNS applies when the target is an ELB.Free-form metadata (up to 255 chars per string; multiple strings per record allowed). Production uses: SPF ("v=spf1 include:amazonses.com ~all"), DKIM selectors (v=DKIM1; k=rsa; p=MIGfMA0GCSq...), DMARC (v=DMARC1; p=reject; rua=mailto:dmarc@example.com), provider ownership verification, and — critically for this tutorial — ACME DNS-01 tokens: _acme-challenge.example.com TXT "LoqXcYV8q5ONbJQxbmR7SCTNo3tiAXDfKYjxAIuNaD6".
Mail routing with a priority (lower wins) and a dot-terminated target: example.com. 300 IN MX 10 inbound-smtp.eu-central-1.amazonaws.com. The target must resolve to A/AAAA — never a CNAME (RFC 2181).
NS records are the delegation fabric: at the apex they advertise the zone's own nameservers; at any other label they create a subdomain delegation. SOA is the zone's metadata record: ns-1024.awsdns-00.org. awsdns-hostmaster.amazon.com. 1 7200 900 1209600 86400 — fields are MNAME (primary NS), RNAME (admin mailbox), serial, refresh, retry, expire, and negative-caching TTL. That last field governs how long resolvers cache NXDOMAIN — the reason a new record can go live in seconds while a deleted record haunts you for hours.
60–300 at least 24–48 h ahead so caches drain. After the cutover, confirm with dig +trace and dig @ns-1024.awsdns-00.org (authoritative) before trusting a recursive resolver's answer.Every LoadBalancer Service and Ingress you create gets a fresh, opaque DNS name (e.g., k8s-prod-app-1234567890.eu-central-1.elb.amazonaws.com). Managing CNAME/Alias records by hand for each of them is unmaintainable. ExternalDNS automates it: it watches Kubernetes sources, diffs the desired state against your Route 53 hosted zone, and executes the minimum ChangeResourceRecordSets API calls to converge.
The sync loop, phase by phase:
Service (type LoadBalancer) and Ingress resources (also Gateway API routes).--txt-owner-id. Ownership prevents two clusters from fighting over the same records.route53:ChangeResourceRecordSets with the target ALB/NLB DNS name as an Alias (for ELB targets) or CNAME.Ownership records look like this in your zone:
external-dns-app.example.com. 300 IN TXT "heritage=external-dns,external-dns/owner=eks-demo,external-dns/resource=ingress/prod/app"
Controlling which records are created is annotation-driven:
metadata:
annotations:
external-dns.alpha.kubernetes.io/hostname: app.example.com
external-dns.alpha.kubernetes.io/ttl: "300"
external-dns.alpha.kubernetes.io/record-type: A
external-dns.alpha.kubernetes.io/ingress-hostname-source: annotation-only
ExternalDNS must be able to mutate Route 53, but it must not carry cluster-admin AWS credentials. The safe pattern on EKS is IAM Roles for Service Accounts (IRSA): your cluster exposes an OIDC issuer, an IAM role's trust policy allows the sts:AssumeRoleWithWebIdentity principal, and the pod's ServiceAccount is annotated with eks.amazonaws.com/role-arn. The kubelet injects a projected OIDC token; the AWS SDK exchanges it for role credentials. No long-lived keys anywhere.
Trust policy for the ExternalDNS role:
{
"Version": "2012-10-17",
"Statement": [
{
"Effect": "Allow",
"Principal": {
"Federated": "arn:aws:iam::123456789012:oidc-provider/oidc.eks.eu-central-1.amazonaws.com/id/9F8A7B6C5D4E3F2A1B0C"
},
"Action": "sts:AssumeRoleWithWebIdentity",
"Condition": {
"StringEquals": {
"oidc.eks.eu-central-1.amazonaws.com/id/9F8A7B6C5D4E3F2A1B0C:aud": "sts.amazonaws.com",
"oidc.eks.eu-central-1.amazonaws.com/id/9F8A7B6C5D4E3F2A1B0C:sub": "system:serviceaccount:external-dns:external-dns"
}
}
}
]
}
Scoped Route 53 policy — write access confined to the one zone, list access required for zone discovery:
{
"Version": "2012-10-17",
"Statement": [
{
"Effect": "Allow",
"Action": ["route53:ChangeResourceRecordSets", "route53:ListResourceRecordSets"],
"Resource": [
"arn:aws:route53:::hostedzone/Z0123456789ABCDEFGHIJ",
"arn:aws:route53:::change/*"
]
},
{
"Effect": "Allow",
"Action": ["route53:ListHostedZones", "route53:ListHostedZonesByName"],
"Resource": "*"
}
]
}
ServiceAccount and Deployment:
apiVersion: v1
kind: ServiceAccount
metadata:
name: external-dns
namespace: external-dns
annotations:
eks.amazonaws.com/role-arn: arn:aws:iam::123456789012:role/external-dns-route53
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: external-dns
namespace: external-dns
spec:
replicas: 1
selector:
matchLabels: { app: external-dns }
template:
metadata:
labels: { app: external-dns }
spec:
serviceAccountName: external-dns
containers:
- name: external-dns
image: registry.k8s.io/external-dns/external-dns:v0.15.1
args:
- --provider=aws
- --registry=txt
- --txt-owner-id=eks-demo
- --policy=upsert-only
- --source=ingress
- --source=service
- --domain-filter=example.com
- --log-level=info
--policy=upsert-only never deletes records, so a misconfigured rollout cannot rip out production DNS. Use it until you are confident, then evaluate --policy=sync for full garbage collection. Pair --domain-filter=example.com with the annotation source so one cluster can never touch another team's zones.Cert-Manager automates the ACME protocol against Let's Encrypt: issuance, renewal, and revocation, all expressed as declarative CRDs. The relevant object model:
| CRD | Scope | Responsibility |
|---|---|---|
Issuer | Namespace | ACME account config (server, email, credentials) for one namespace |
ClusterIssuer | Cluster | Same, but cluster-wide — referenced by cert-manager.io/cluster-issuer annotations from any namespace |
Certificate | Namespace | Desired certificate: DNS names, secret name, issuer reference, duration/renewBefore |
CertificateRequest | Namespace | One-shot signing request with CSR — the issuer-specific unit of work |
Order | Namespace | An ACME order lifecycle (one or more authorizations) against the ACME directory |
Challenge | Namespace | A single authorization attempt with the chosen solver (http01/dns01) |
Issuance pipeline (DNS-01 flavor):
Ingress with cert-manager.io/cluster-issuer and a tls block → the ingress-shim controller synthesizes a Certificate.CertificateRequest (CSR).Order against https://acme-v02.api.letsencrypt.org/directory.Challenge is created; the solver performs the verification (section 3).Secret (keys tls.crt, tls.key, ca.crt), which the Ingress controller picks up automatically.Let's Encrypt certificates are valid 90 days; cert-manager renews when the certificate is within renewBefore (default 30 days) of expiry, with jitter to spread load. Track fleet-wide state with kubectl get certificate -A -o wide and per-issue status via kubectl describe order,challenge -n prod.
Let's Encrypt issues a token and expects it served at http://app.example.com/.well-known/acme-challenge/LoqXcYV8q5ONbJQxbmR7SCTNo3tiAXDfKYjxAIuNaD6 over port 80. The response body must be the token concatenated with your account key thumbprint (key authorization). Cert-Manager solves this by creating a temporary ingress (and backing service) on the fly that serves the token, then deleting it after validation.
/.well-known/acme-challenge/, no CDN that rewrites the path.*.example.com.Let's Encrypt gives you a key authorization and expects it published as a TXT record at _acme-challenge.<name>. For app.example.com the record is _acme-challenge.app.example.com; for a wildcard *.example.com the record goes at _acme-challenge.example.com (the wildcard label is stripped). The CA then queries DNS directly and compares the value.
route53 solver calls ChangeResourceRecordSets with the TXT value, then waits for propagation (default --dns01-propagation-wait 120 s) while polling the authoritative nameservers._acme-challenge TXT records simultaneously (one per authorization).| Dimension | HTTP-01 | DNS-01 |
|---|---|---|
| Mechanism | CA fetches token from http://<host>/.well-known/acme-challenge/<token> over port 80 | CA queries TXT at _acme-challenge.<name> via normal DNS resolution |
| Route 53 permissions | None — no DNS mutation, works with a bare public ingress | route53:ChangeResourceRecordSets + ListResourceRecordSets scoped to the hosted zone; IRSA role on the cert-manager pod |
| Internet visibility | Requires public reachability on port 80; breaks on private VPCs, hairpin NAT, WAF/CDN path rewriting | None inbound — only egress to Let's Encrypt and the Route 53 API; works in fully private clusters |
| Wildcard support | No | Yes (*.example.com → TXT at _acme-challenge.example.com) |
| Pros | Zero IAM footprint; fastest bootstrap; no propagation wait; no DNS provider credentials in cluster | Works everywhere (private VPCs, wildcards); no open inbound port; validation invisible to users; identical flow for every name |
| Cons | Port 80 must be open per host; no wildcards; flaky behind aggressive WAFs/redirect chains | DNS provider credentials must live in the cluster; propagation delay (TTL + polling); broader blast radius if the role leaks |
--dns01-recursive-nameservers).Everything below assumes an EKS cluster in eu-central-1 (account 123456789012), a public hosted zone example.com (Zone ID Z0123456789ABCDEFGHIJ), and an NGINX Ingress Controller installed. Cert-manager is installed with Helm:
helm repo add jetstack https://charts.jetstack.io
helm install cert-manager jetstack/cert-manager \
--namespace cert-manager --create-namespace \
--set crds.enabled=true \
--set serviceAccount.annotations."eks\.amazonaws\.com/role-arn"=arn:aws:iam::123456789012:role/cert-manager-route53
IRSA role cert-manager-route53 with the same OIDC trust pattern as ExternalDNS, restricted to system:serviceaccount:cert-manager:cert-manager:
{
"Version": "2012-10-17",
"Statement": [
{
"Effect": "Allow",
"Action": ["route53:GetChange", "route53:ListResourceRecordSets"],
"Resource": "arn:aws:route53:::change/*"
},
{
"Effect": "Allow",
"Action": ["route53:ChangeResourceRecordSets", "route53:ListResourceRecordSets"],
"Resource": "arn:aws:route53:::hostedzone/Z0123456789ABCDEFGHIJ"
},
{
"Effect": "Allow",
"Action": ["route53:ListHostedZones", "route53:ListHostedZonesByName"],
"Resource": "*"
}
]
}
The ChangeResourceRecordSets grant is the dangerous one — scope it to the single zone so a compromised cert-manager cannot touch unrelated DNS.
apiVersion: cert-manager.io/v1
kind: ClusterIssuer
metadata:
name: letsencrypt-prod
spec:
acme:
server: https://acme-v02.api.letsencrypt.org/directory
email: ops@example.com
privateKeySecretRef:
name: letsencrypt-prod-account-key
solvers:
- dns01:
route53:
hostedZoneID: Z0123456789ABCDEFGHIJ
region: eu-central-1
The solver uses the pod's ambient credentials (IRSA) — no access keys in YAML. If the cluster is not running inside AWS, or you need cross-account issuance, add accessKeyIDSecretRef/role instead; prefer IRSA whenever you can.
apiVersion: cert-manager.io/v1
kind: ClusterIssuer
metadata:
name: letsencrypt-http
spec:
acme:
server: https://acme-v02.api.letsencrypt.org/directory
email: ops@example.com
privateKeySecretRef:
name: letsencrypt-http-account-key
solvers:
- http01:
ingress:
class: nginx
Requirements: the NGINX controller must be publicly reachable on port 80, and the temporary ingress it creates for /.well-known/acme-challenge/ must not be filtered by network policies. No IAM is needed.
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
name: app
namespace: prod
annotations:
kubernetes.io/ingress.class: nginx
external-dns.alpha.kubernetes.io/hostname: app.example.com
external-dns.alpha.kubernetes.io/ttl: "300"
cert-manager.io/cluster-issuer: letsencrypt-prod
spec:
tls:
- hosts:
- app.example.com
secretName: app-example-tls
rules:
- host: app.example.com
http:
paths:
- path: /
pathType: Prefix
backend:
service:
name: app
port:
number: 80
For a wildcard, swap the annotation to external-dns.alpha.kubernetes.io/hostname: "*.example.com" and the TLS block to hosts: ["*.example.com"] with secretName: wildcard-example-tls — and make sure the issuer is the DNS-01 one, because wildcards never issue over HTTP-01.
external-dns.alpha.kubernetes.io/hostname annotation wins over the Ingress host field. Keep them identical, or ExternalDNS will churn records every reconcile.# DNS record appeared
kubectl get ingress app -n prod -o jsonpath='{.status.loadBalancer.ingress[0].hostname}'
dig +short app.example.com
# Certificate state
kubectl get certificate -n prod app-example-tls -o wide
kubectl get order,challenge -n prod
# Secret populated with a full chain
kubectl get secret app-example-tls -n prod -o jsonpath='{.data.tls\.crt}' | base64 -d | openssl x509 -noout -subject -dates
dig A app.example.com
dig AAAA app.example.com
dig TXT _acme-challenge.example.com
dig NS example.com
dig MX example.com
dig SOA example.com
Expected output for dig A app.example.com:
; <<>> DiG 9.18.19 <<>> A app.example.com
;; global options: +cmd
;; Got answer:
;; ->>HEADER<<- opcode: QUERY, status: NOERROR, id: 31415
;; flags: qr rd ra; QUERY: 1, ANSWER: 1, AUTHORITY: 0, ADDITIONAL: 1
;; QUESTION SECTION:
;app.example.com. IN A
;; ANSWER SECTION:
app.example.com. 300 IN A 203.0.113.42
;; Query time: 12 msec
;; SERVER: 127.0.0.53#53(127.0.0.53) (UDP)
;; WHEN: Wed Aug 20 09:41:23 UTC 2026
;; MSG SIZE rcvd: 62
dig +trace example.com
; <<>> DiG 9.18.19 <<>> +trace example.com
;; global options: +cmd
. 518400 IN NS a.root-servers.net.
;; Received 525 bytes from 198.41.0.4#53(a.root-servers.net) in 23 ms
com. 172800 IN NS a.gtld-servers.net.
;; Received 758 bytes from 192.5.6.30#53(a.gtld-servers.net) in 71 ms
example.com. 172800 IN NS ns-1024.awsdns-00.org.
example.com. 172800 IN NS ns-2048.awsdns-01.com.
example.com. 172800 IN NS ns-3072.awsdns-02.co.uk.
example.com. 172800 IN NS ns-4096.awsdns-03.net.
;; Received 760 bytes from 192.12.94.30#53(a.gtld-servers.net) in 88 ms
app.example.com. 300 IN A 203.0.113.42
;; Received 62 bytes from 205.251.192.4#53(ns-1024.awsdns-00.org) in 41 ms
Each section shows a hop: root → TLD → your zone's NS delegation → final authoritative answer. If the last NS block does not list awsdns nameservers, your delegation points somewhere else — that is your first clue.
# Ask a specific authoritative nameserver; +norecurse forbids it from recursing for us
dig @ns-1024.awsdns-00.org example.com A +norecurse +noall +answer
# Compare with what a public recursive resolver (8.8.8.8) still has cached
dig @8.8.8.8 app.example.com A
# Short, scriptable output
dig +short app.example.com
dig +short TXT _acme-challenge.example.com
dig +short NS example.com
$ dig @ns-1024.awsdns-00.org example.com A +norecurse +noall +answer
example.com. 300 IN A 203.0.113.42
$ dig +short app.example.com
203.0.113.42
$ dig +short TXT _acme-challenge.example.com
"LoqXcYV8q5ONbJQxbmR7SCTNo3tiAXDfKYjxAIuNaD6"
Handy flags: +short (answer only), +trace (walk delegation), +norecurse (authoritative-only answer), +noall +answer (strip comments/sections), +time=2 +tries=2 (query timeout/retries), -t TYPE (query type).
# Default recursive query
nslookup app.example.com
# Specific record type
nslookup -type=ns example.com
nslookup -type=txt _acme-challenge.example.com
# Query a specific DNS server (bypass local cache)
nslookup app.example.com 8.8.8.8
nslookup -type=soa example.com ns-1024.awsdns-00.org
$ nslookup app.example.com
Server: 127.0.0.53
Address: 127.0.0.53#53
Non-authoritative answer:
Name: app.example.com
Address: 203.0.113.42
Interactive mode — change server and type mid-session:
$ nslookup
> server 8.8.8.8
Default server: 8.8.8.8
Address: 8.8.8.8#53
> set type=ns
> example.com
Server: 8.8.8.8
Address: 8.8.8.8#53
Non-authoritative answer:
example.com nameserver = ns-1024.awsdns-00.org.
example.com nameserver = ns-2048.awsdns-01.com.
example.com nameserver = ns-3072.awsdns-02.co.uk.
example.com nameserver = ns-4096.awsdns-03.net.
> set type=txt
> _acme-challenge.example.com
Server: 8.8.8.8
Address: 8.8.8.8#53
_acme-challenge.example.com text = "LoqXcYV8q5ONbJQxbmR7SCTNo3tiAXDfKYjxAIuNaD6"
> exit
kubectl apply -f ingress-app.yamlkubectl logs deploy/external-dns -n external-dns --tail=100 → expect Creating record: app.example.com A ..., then dig +short app.example.com returns the ALB DNS name or IP.kubectl get certificate,order,challenge -n prod -w; during DNS-01, confirm the TXT is visible authoritatively (propagation) while the recursive cache may lag: dig @ns-1024.awsdns-00.org TXT _acme-challenge.app.example.com.kubectl describe certificate app-example-tls -n prod → Ready=True; verify over the wire: echo | openssl s_client -connect app.example.com:443 -servername app.example.com 2>/dev/null | openssl x509 -noout -subject -issuer -dates.kubectl get certificate -A -o custom-columns=NAME:.metadata.name,READY:.status.conditions[0].status,NEXT:.status.renewalTime.| Symptom | Cause | Fix |
|---|---|---|
dig still returns the old IP after a change | TTL + resolver caching; you asked a recursive server | Query authoritative directly: dig @ns-1024.awsdns-00.org example.com A +norecurse |
Certificate stuck on Waiting for DNS-01 challenge | IAM/Route 53 policy missing, wrong hosted zone ID, or TXT not propagating | Check kubectl describe challenge, then the dig @ns-1024.awsdns-00.org TXT _acme-challenge.example.com path |
| HTTP-01 challenge 404s | Port 80 blocked, WAF filtering, or ingress class mismatch | Confirm curl http://app.example.com/.well-known/acme-challenge/ reachable and unrewritten |
| ExternalDNS logs but creates nothing | Missing hostname annotation, or --domain-filter excludes the name | Add the annotation; verify zone discovery in logs; run with --log-level=debug |
| Ingress serves the old certificate after rotation | NGINX cached the secret; secret name reused across hosts | Use a unique secretName per host; restart the controller or wait for its resync |