DNS Troubleshooting Guide: From resolv.conf to Kubernetes - Resolving Intermittent Failures and Latency
A comprehensive guide to diagnosing DNS resolution latency and failures, covering resolv.conf configuration, ndots and search domain pitfalls, caching layers, Kubernetes DNS specifics, and practical troubleshooting steps with command-line tools.
Problem Background
DNS resolution is an often overlooked "hidden bomb" that can take down entire services. Common symptoms include applications occasionally hanging for tens of seconds then failing (retry succeeds), upstream DNS changes causing intermittent resolution, short names resolving to wrong hosts due to search domains, container resolution delays from ndots:5, and sudden latency spikes after upgrades.
Applicable Scenarios
Application domain resolution slow, timeout, intermittent, or occasional failures
Behavior changes after modifying resolv.conf, /etc/nsswitch.conf, /etc/hosts, or DNS configs
Internal services using short names ( db, redis) resolving to wrong hosts or high latency
Container/K8s Pod resolution behavior differs from bare metal
Need to distinguish DNS issues from network/TLS/connection issues
Need to unify DNS config across machines with validation and rollback
Prerequisite: network layer reachable (ping target IP works). If IP unreachable, fix connectivity first.
Core Knowledge
Resolution Flow
Application calls getaddrinfo() or similar
Query order determined by
/etc/nsswitch.conf hosts:line (commonly files dns)
DNS queries follow
/etc/resolv.conf nameserverentries sequentially
Query name handling depends on ndots and search (direct FQDN vs search domain suffixes)
Results may be cached by nscd, systemd-resolved, or application caches
Each nameserver has timeout (default 5s) and attempts (default 2)
Key resolv.conf Fields
nameserver 10.0.0.53 # up to 3, tried in order (or rotate with options rotate)
search example.com sub.example.com # multiple search domains
options timeout:1 attempts:2 ndots:1 rotate nameserver: max 3, tried sequentially unless rotate set search / domain: append suffixes to non-FQDN names ndots: dots threshold to treat as FQDN (default 1) timeout: per-nameserver query timeout seconds (default 5) attempts: retries per nameserver (default 2) rotate: round-robin across nameservers
ndots Deep Dive
ndots:Nmeans names with ≥N dots are queried as FQDN directly. Default ndots:1. With search example.com: host.example.com (1 dot) → direct query host (0 dots) → tries host.example.com then host Kubernetes often sets ndots:5, causing short names and even single-dot names to attempt multiple search suffixes, each adding latency and potential timeouts. This is the #1 suspect for container resolution slowness.
search Domain Pitfalls
Search suffixes can collide with external domains (e.g., db → db.prod.example.com conflicting with public same name), causing "resolves but to wrong host". Unresolvable search domains waste timeout cycles per query.
/etc/hosts & nsswitch.conf
/etc/hosts: static mappings, highest priority if files first in hosts: line
/etc/nsswitch.conf hosts:determines lookup order (e.g., files dns)
Check /etc/hosts for pollution when results mismatch expectations
Cache Layers
nscd(legacy) / systemd-resolved (modern): cache resolution results
Application-level caches (Java, Go, Node each have own strategies)
Config changes not reflecting? Suspect cache first
Upstream & Recursion
Configured nameservers are usually recursive resolvers (VPC DNS, cloud DNS, 8.8.8.8). Slowness may be upstream or path to upstream. Use dig +trace to expand recursive path and see per-hop latency.
Overall Troubleshooting Approach
1. Confirm DNS issue (compare IP direct connect)
2. Check local resolver config (resolv.conf / nsswitch.conf / hosts)
3. Use dig to reproduce and separate local config from upstream
4. Verify ndots/search rewrite effects
5. Check timeout/retry/cache layers (timeout, attempts, nscd/resolved)
6. Verify time (NTP) and multiple nameserver order/failures
7. Fix + validate
8. Rollback plan + prevention (monitoring & cache strategy)Practical Troubleshooting Steps
Step 1: Confirm DNS vs Connectivity
# Test connectivity directly with IP
ping -c 3 <targetIP>
curl -v --connect-timeout 3 http://<targetIP>/
# Measure resolution time
time getent hosts <domain>
time nslookup <domain> getent hostsslow/fails → DNS issue, continue getent hosts fast, IP curl works → issue may be in app's resolver/cache/proxy
IP unreachable → network issue, not DNS
Note: getent hosts uses glibc stack (hosts+DNS), dig queries DNS directly bypassing /etc/hosts. Different behaviors help isolate layer.
Step 2: Check Local Config
cat /etc/resolv.conf
cat /etc/nsswitch.conf | grep hosts
cat /etc/hostsAre nameserver IPs reachable and appropriate?
Is hosts: order changed unexpectedly?
Any problematic entries in /etc/hosts?
If resolv.conf managed by DHCP/systemd-resolved, edit source not the symlink
Step 3: Use dig to Isolate Local Stack vs Upstream
# Direct query to specific server, bypass search
dig @10.0.0.53 <FQDN>
# Measure latency
dig @10.0.0.53 <FQDN> +time=2 +tries=1
# See search expansion for short name
dig <shortname> +search +time=2
# Trace recursive path
dig +trace <FQDN> dig @server FQDNfast + correct → upstream OK, problem local (ndots/search/order/cache) dig @server FQDN slow/timeout → upstream or path issue dig shortname +search shows actual expanded queries +trace reveals per-hop latency (root, TLD, authoritative)
Step 4: Verify ndots & search Rewrites
cat /etc/resolv.conf | grep -E 'ndots|search|options'
getent hosts <shortname> || true
dig <shortname> +search +shortWith ndots:5 and search a b c, short name tries name.a, name.b, name.c, name sequentially. Each adds latency.
Optimization: lower ndots to 1 or 0 for FQDN-heavy apps, trim search domains, or use trailing dot ( name.example.com.) for absolute names.
Step 5: Timeouts, Retries & Cache
grep -E 'timeout|attempts|rotate' /etc/resolv.conf
systemctl status nscd systemd-resolved 2>/dev/null
ls -l /etc/resolv.conf
readlink -f /etc/resolv.conf
resolvectl status 2>/dev/null | headDefault timeout:5 attempts:2 can yield tens of seconds worst-case, amplified by search variants
Config change no effect? Clear cache: systemctl restart nscd / systemctl restart systemd-resolved / app restart
If systemd-resolved manages resolv.conf (symlink to /run/systemd/resolve/...), edit /etc/systemd/resolved.conf not the symlink
Step 6: NTP & Multiple Nameservers
timedatectl
timedatectl status | grep -E 'synchronized|NTP'
grep nameserver /etc/resolv.confLarge time skew affects DNSSEC/TLS/recursive policies → sync time first
First nameserver unreachable/slow drags all queries (unless rotate). Use options rotate or put reliable server first
Step 7: Fix & Validate
Fix resolv.conf → retest with dig @server FQDN +time=2 +tries=1 and getent hosts FQDN If systemd-resolved: edit /etc/systemd/resolved.conf ( DNS=, Domains=) then systemctl restart systemd-resolved Adjust ndots /trim search for short-name scenarios
Validation must cover: getent hosts (glibc), dig (upstream), actual app requests — all fast/normal.
Common Commands Cheatsheet
# Resolution & timing
dig @server name
dig name +short
dig name +trace
getent hosts name
nslookup name
host name
# Config
cat /etc/resolv.conf
cat /etc/nsswitch.conf | grep hosts
cat /etc/hosts
ls -l /etc/resolv.conf
# Cache & services
systemctl status nscd systemd-resolved
systemctl restart nscd systemd-resolved
resolvectl status
# Time
timedatectl status
# Packet capture
tcpdump -i eth0 -n port 53 -c 50Root Cause Deep Dives
7.1 Nameserver Unreachable / Wrong Order
cat /etc/resolv.conf | grep nameserver
for ns in $(awk '/nameserver/{print $2}' /etc/resolv.conf); do
echo "== $ns =="; timeout 2 nc -vz "$ns" 53 2>&1 || echo "$ns:53 unreachable"
doneFix: move reachable/fast servers first, add rotate, or remove unreachable.
7.2 High ndots Latency
cat /etc/resolv.conf | grep ndots
# or
resolvectl status
time dig shortname +search +time=2 | grep -E 'time|status'Fix: lower ndots, trim search, or use trailing dot. Validate with time getent hosts name.example.com. vs time getent hosts name.example.com.
7.3 Search Domain Collision
getent hosts <shortname>
dig <shortname> +searchFix: remove extra search domains, use FQDN, or pin in /etc/hosts (but record it to avoid migration issues).
7.4 Cache Staleness
systemctl restart nscd
systemctl restart systemd-resolved
getent hosts <FQDN>7.5 Time Skew
timedatectl
timedatectl set-ntp on
# or chrony/ntpdate7.6 Application-Layer Caches (Java, Go, Node)
Java: networkaddress.cache.ttl JVM property. Go: net.Dialer and LOOKUP env. Node: process-internal DNS cache. If system dig works but app fails, check app cache/TTL/restart.
7.7 DNSSEC Validation Failures
dig @server FQDN
# SERVFAIL (not NXDOMAIN) suggests DNSSEC issueVerify with DNSSEC disabled or public recursive; coordinate with upstream.
Configuration Examples
8.1 Internal FQDN-Heavy Apps
nameserver 10.0.0.53
nameserver 10.0.1.53
search example.com
options timeout:1 attempts:2 ndots:1 ndots:1lets single-dot names query directly as FQDN; short names use search.
8.2 Short-Name Heavy Environments
nameserver 10.0.0.53
search prod.example.com staging.example.com
options timeout:2 attempts:2 ndots:1 rotate rotatebalances across nameservers.
8.3 Kubernetes/Container Notes
Pod dnsConfig may override ndots (often default ndots:5). Adjust via spec.dnsConfig.options (e.g., ndots:1), but evaluate impact on other resolutions. Example:
dnsPolicy: ClusterFirst
dnsConfig:
options:
- name: ndots
value: "1"
- name: single-request-reopenNote: single-request-reopen behavior varies by resolver (glibc/musl/app); test before adopting.
Log & Metrics Observation
DNS often lacks dedicated logs; rely on active probes and packet capture. Most reliable: tcpdump -i any -n 'udp port 53' -c 100 and tcp port 53. Observe query destination, per-query latency, retries, failures ( [.] root marker), SERVFAIL.
If systemd-resolved: resolvectl statistics shows hit rate and latency.
Baseline: LAN resolution typically milliseconds; cross-network/public recursive tens to hundreds of ms. Alert on sustained >hundreds ms, timeouts, SERVFAIL spikes, duplicate queries.
Decision Tree (Troubleshooting Path)
Domain resolution slow/fail?
A. IP direct connect?
No → not DNS, check network
Yes → next
B. dig @server FQDN OK?
No → upstream/path issue
Yes → next
C. getent hosts FQDN OK?
No → local config/stack/cache/time
Yes → next
D. Short name vs FQDN same?
Different → focus ndots/search
Same → check cache/timeout/multi-nameserver
E. Packet capture latency/failure where?
Local early hops → config/rewrite
Upstream no response → upstream/firewall UDP53Risk Warnings
Modifying /etc/resolv.conf is system config change; backup first
Manual edits may be overwritten by DHCP/systemd-resolved; use proper channels
Removing search domains may break services relying on short names
Lowering ndots may cause previously search-handled short names to fail as bare queries rotate changes query behavior; validate before production
Restarting nscd/systemd-resolved clears global cache — brief jitter, do off-peak
Pinning in /etc/hosts bypasses DNS; causes "migrated but still resolves old IP" issues — document
Always backup before delete/overwrite; provide rollback content
Production changes: backup, canary, rollback, record
Validation Checklist
Backup exists:
ls -l /etc/resolv.conf.bak* dig @server FQDNreturns correct RCODE quickly getent hosts FQDN returns fast
App actual connections normal
Packet capture confirms query count/latency as expected (short names no extra variants)
Multiple nameserver order/rotate behaves as expected
One cycle post-change: no alerts, no new errors
Rollback Procedure
# Backup
cp /etc/resolv.conf /etc/resolv.conf.bak.$(date +%F)
# Rollback (restore pre-change file)
cp /etc/resolv.conf.bak.$(date +%F) /etc/resolv.conf
# If systemd-resolved mode: restore resolved.conf then restart
# Record original resolve config, rollback then systemctl restart systemd-resolvedIf restart nscd caused jitter: wait for self-heal or restart services sequentially
Post-rollback verify getent hosts and app recovery
Production Best Practices
Change window: low traffic, notify affected teams
Canary single node then batch
Unified config via config management to prevent drift
Multi-tenant/cluster: watch interaction with app-layer caches and Pod overrides
Monitor "resolution success rate + latency"; alert on anomalies
Document: owner, upstream DNS, dependencies, rollback method
Include NTP sync in monitoring; time drift amplifies DNS issues
Prevention: DNS in Routine Ops
Baseline: record normal resolution latency/success rate for alert thresholds
Metrics: DNS query success rate, avg latency, SERVFAIL count (per actual exporter)
Quarterly audit: resolv.conf consistency, prune stale search domains, check expired /etc/hosts entries
Change log: every DNS/nameserver/ ndots change recorded
Drill: periodic "kill nameserver + rollback" exercise for critical domains
Full Case Study: Container Intermittent Timeout → ndots
Symptom: Microservice calling third-party occasionally hangs 10+ seconds, retry succeeds.
Troubleshoot: IP direct connect OK. dig @vpc-dns FQDN +time=2 fast. But getent hosts for internal short name slow. /etc/resolv.conf shows options ndots:5 search prod.example.com staging.example.com. Short name tries multiple search variants, each failed path adds wait.
Fix: For "mostly FQDN" service, lower ndots to 1 or use FQDN with trailing dot. Canary single node then batch.
Validate: time getent hosts FQDN and packet capture confirm query count drops to 1, latency back to milliseconds. Rollback: record original resolv /override config, restore on anomaly.
Root cause: ndots:5 × multiple search domains × timeout stacking. Lesson: check ndots / search rewrite impact before staring at nameservers.
dig Subcommands Deep Dive
17.1 Basic Query & RCODE
dig @10.0.0.53 example.com ARead status ( NOERROR / NXDOMAIN / SERVFAIL), ANSWER TTL, Query time (most direct "how long").
17.2 Force No Cache, No Search
dig +norecurse example.com # non-recursive, authoritative view
dig example.com +short # just results
dig example.com +noall +answer +norecurseon recursive server returns NOERROR but no ra flag; shows if authoritative side has issue.
17.3 Trace Recursive Path
dig example.com +traceShows each hop from root to authoritative with Query time. Slow hop = latency source.
17.4 Query Type & TCP
dig @server example.com CNAME
dig @server +tcp example.comUDP default; some envs filter UDP or need TCP for large responses. +tcp tests TCP path.
17.5 Timeout & Retry Control
dig @server example.com +time=2 +tries=1Simulates "fail fast" to see if network issues cause long hangs.
17.6 Batch Queries
for n in db redis cache; do echo "== $n =="; dig $n +short; doneQuick scan which short names slow/fail.
FAQ
19.1 Why dig fast but getent hosts slow?
digqueries specified nameserver directly, may skip search expansion and some cache layers. getent hosts uses system stack: /etc/hosts, nsswitch, ndots/search, /etc/resolv.conf. Discrepancy points to problematic layer.
19.2 How ndots affects speed?
ndotsdecides dot threshold for FQDN treatment. Below threshold, search suffixes tried sequentially; each failed variant consumes attempt/wait. ndots:5 makes short names try many variants. Fix: caller uses name. (trailing dot) or ndots:1.
19.3 resolv.conf change not effective?
Likely managed by systemd-resolved or DHCP. readlink -f /etc/resolv.conf shows symlink; edit resolved config source, not symlink target.
19.4 Same domain sometimes works sometimes not?
One nameserver unreachable, timeout too small, or load spread to slow backend. Loop dig multiple times to see failure rate; add rotate or reorder nameservers.
19.5 Should I pin internal domains in /etc/hosts?
Can but cautiously. /etc/hosts highest priority, bypasses DNS. Solves "can't resolve internal" but creates "IP changed but still resolves old" risk. If used: document, limit scope, monitor for cleanup.
19.6 Does time skew affect resolution?
Yes. Large skew breaks DNSSEC validation and some recursive policies. timedatectl check; run NTP/chrony.
19.7 App cached DNS?
Java/Go/Node have own DNS caches or resolver strategies. /etc/hosts and resolv.conf changes not picked up is normal process-internal cache. Fix: adjust app TTL or restart app; not in system config.
DNS Change Compliance Template
Change Title: app-01 adjust DNS search domains and ndots
Reason: Container short-name resolution latency from ndots:5 + multiple search
Machines: app cluster (canary app-01 → two batches)
Pre-check: dig/getent baseline latency; confirm /etc/resolv.conf ownership (resolved?)
Action: backup original → apply; canary one → validate getent/dig/app → batch
Files: /etc/resolv.conf (or systemd-resolved source)
Validation: getent latency drop + dig query count down + app no errors
Rollback: cp backup restore + systemctl restart nscd/resolved; on anomaly revert
Impact: all name resolution on machine; declare affected services
Approver/Executor: ____Monitoring Probe Script
Read-only probe /usr/local/bin/dns-health.sh collects key FQDN resolution latency and status, cron-able, monitoring-ingestible.
#!/usr/bin/env bash
# DNS health probe (read-only)
NAME=${1:-www.example.com}
SERVERS=(${@:2})
[ ${#SERVERS[@]} -eq 0 ] && SERVERS=("10.0.0.53" "119.29.29.29")
for ns in ${SERVERS[@]}; do
out=$(dig @"$ns" "$NAME" +time=2 +tries=1 2>/dev/null)
st=$(printf '%s
' "$out" | awk '/status:/{print $4}')
qt=$(printf '%s
' "$out" | awk '/Query time/{print $4}')
echo "$(date +%F\ %T) $ns $NAME status=$st qtime=${qt}ms"
done
# System stack perspective
start=$(date +%s%3N)
getent hosts "$NAME" >/dev/null 2>&1; rc=$?
end=$(date +%s%3N)
echo "system-stack name=$NAME rc=$rc ms=$((end-start))"Usage:
sudo bash /usr/local/bin/dns-health.sh www.example.com 10.0.0.53. Non-intrusive; log to cron or push metrics. Alert on non- NOERROR status or qtime exceeding baseline.
Rollback Nuances: Preventing Recurrence
Manual /etc/resolv.conf edit later overwritten by DHCP/resolved → rollback must align with source, else re-overwritten
systemd-resolved rollback = restore original resolved.conf then restart
App-layer TTL changes need rollback to original config /etc/hosts pinned records rollback = delete lines
Rollback isn't just "revert values in 3 seconds"; must ensure source won't re-apply. Often overlooked, causes re-recurrence.
Minimal Monitoring & Alerting Loop
Collect: scheduled dns-health.sh → push qtime / status / rc to logs/metrics
Metrics: resolution success rate, avg latency, SERVFAIL / NXDOMAIN counts (per collector)
Thresholds: sustained 0% success or latency > baseline N× → alert; tune per business cycle, no absolute hard numbers
Linkage: alert triggers "dig → getent → tcpdump" mainline auto/manual reproduction
IPv6 & DNS: Don't Only Debug IPv4
Many "slow/weird" resolutions involve IPv6 interference. Check app/system v6 resolution path:
ip -6 addr
cat /etc/gai.conf | grep -v '^#' | head
dig @10.0.0.53 www.example.com AAAA +time=2
getent ahosts www.example.com | headCommon pitfalls: getent ahosts returns mixed A/AAAA; app may stall waiting for non-existent AAAA ("AAAA stall") /etc/gai.conf address selection (RFC 6724) reorders IPv4 vs IPv6 preference
IPv6 firewall ( ip6tables /nft ip6) unconfigured → v6 connections blocked/loopback
Logic: App prefers AAAA but no v6 connectivity → extra wait. Adjust gai.conf order or force v4 in app. Validate v4/v6 separately with ping6 / curl -6/-4. Backup gai.conf before changes; canary validate.
DoH/DoT & Forwarders Reminder
Modern stacks use dnscrypt / DoH / DoT, sslh, unbound, dnsmasq. If resolv.conf nameserver is 127.0.0.53 / 127.0.0.1, real upstream is in forwarder config.
grep nameserver /etc/resolv.conf
cat /etc/systemd/resolved.conf 2>/dev/null
cat /etc/unbound/unbound.conf 2>/dev/null | grep -iE 'forward|tcp|tls'If nameserver is localhost stub, editing resolv.conf nameserver may be ineffective; edit forwarder component. Explains "I changed resolv.conf but nothing changed".
Quantitative: qtime, Hit Count, Failure Rate Together
Single slow query ≠ incident. Quantify. Collection pseudocode:
for i in $(seq 1 20); do
dig @10.0.0.53 www.example.com +time=1 +tries=1 2>/dev/null
done | grep -c 'NOERROR' # success countLogic: Success rate significantly <100% or latency exceeds baseline → formal investigation. Transient single slow query not incident; look at trends and multiple metrics (qtime, SERVFAIL count, reachability) jointly.
Cross-Machine DNS Config Consistency
Many machines → DNS drift causes "this works, that doesn't". Use config management:
- name: Deploy resolv.conf
ansible.builtin.copy:
content: |
nameserver 10.0.0.53
search example.com
options timeout:2 attempts:2 ndots:1
dest: /etc/resolv.conf
notify: restart resolv-managerIf all resolv.conf consistent but differences persist → network/upstream/regional diff, don't blame DNS alone. Rollback = restore old resolv.conf and config.
One-Page DNS Troubleshooting Checklist
☐ 1. IP direct connect OK? (ping/curl target IP)
☐ 2. dig @server FQDN fast and NOERROR?
☐ 3. getent hosts FQDN fast?
☐ 4. resolv.conf nameservers reachable and first?
☐ 5. systemd-resolved/DHCP managed? (readlink)
☐ 6. ndots too high? search domains excessive?
☐ 7. Short name expanded to wrong domain?
☐ 8. After config change, nscd/resolved cache cleared?
☐ 9. App-layer independent cache/TTL?
☐ 10. Time sync (NTP) OK?
☐ 11. IPv6 interference?
☐ 12. Forwarder layer (unbound/dnsmasq) intercepting?
☐ 13. Packet capture (tcpdump port 53) confirms query path
☐ 14. Baseline/rollback readyIf checklist exhausted and root cause not found, issue likely further upstream in recursive/authoritative chain or requires coordination. This checklist maps 1:1 with article sections — condensed execution version.
Field Case Studies
Case 1: Resolved IP but Connection Hangs at Trying
Container curl -v https://www.example.com/ shows "Host resolved" then "Trying 203.0.113.20:443..." timeout. DNS already succeeded. Need to trace SYN egress: container → pod NIC → node → firewall/NAT → SYN-ACK.
getent ahostsv4 www.example.com
curl -4v --connect-timeout 5 https://www.example.com/
ip route get 203.0.113.20
sudo tcpdump -ni any 'host 203.0.113.20 and tcp port 443'Only SYN retransmits → check NetworkPolicy, node egress, SNAT, upstream ACL, target service. TLS stall after handshake → verify SNI, cert, proxy, middlebox. Use curl --resolve domain:443:IP https://domain/ to preserve Host/SNI. Measure DNS/TCP/TLS phases with
curl -w '%{time_namelookup} %{time_connect} %{time_appconnect}
'.
Case 2: Pod External Domain Latency Amplified by search Domains
Pod resolv.conf has three K8s search suffixes and ndots:5. App heavily accesses api.example.com. Under glibc, dot count <5 → tries search variants first. If each candidate times out, external domain first-resolution latency amplifies; if quick NXDOMAIN, impact small. Must measure actual queries and latency.
cat /etc/resolv.conf
getent ahosts api.example.com
dig api.example.com. A +time=2 +tries=1
sudo tcpdump -ni any -vv 'port 53'Validate: adjust dnsConfig.options.ndots or use trailing dot for specific FQDN, compare real app (glibc/Go/Java differ) P95 resolution latency, DNS QPS, cluster short service calls. Don't blindly set ndots:0 — breaks service.namespace internal short names. Modify via workload declaration, canary, not hand-edit container file.
Case 3: dig OK but App Reports Cannot Resolve
dig @10.43.0.10 api.example.comreturns A record, but Python/Java says domain not exist. dig direct query ≠ NSS, /etc/hosts, systemd-resolved, local cache, nor proves app's network namespace matches test shell.
getent hosts api.example.com
sed -n '/^hosts:/p' /etc/nsswitch.conf
cat /etc/hosts
readlink -f /etc/resolv.conf
sudo nsenter -t <pid> -n getent hosts api.example.com
resolvectl query api.example.comDivergence: getent fails, dig works → check NSS order, hosts override, local stub, search domains. System call OK but app fails → app DNS cache, proxy config, runtime independent resolver, container file view. Record current resolution and cache TTL before clearing; mere app restart may mask upstream issue.
Case 4: A Record OK, IPv6 Path Times Out
Domain has A and AAAA; app connects via IPv6 slowly then falls back to IPv4. DNS not failing; issue in address selection and IPv6 routing. Only dig A misses this branch.
dig api.example.com A +short
dig api.example.com AAAA +short
getent ahosts api.example.com
curl -4v --connect-timeout 3 https://api.example.com/
curl -6v --connect-timeout 3 https://api.example.com/Cross-reference ip -6 route, container IPv6 support, egress firewall, proxy AAAA handling, target service listeners. Force IPv4 only as temporary diagnostic/recorded fallback, not delete public AAAA before root cause clear. Retest must cover IPv4/IPv6 success rates and client selection path.
Case 5: Multiple Nameservers + Cache → "Fix Not Effective"
First nameserver intermittent packet loss, second healthy; app shows periodic long tail. Resolvers differ in retry/rotate/cache strategies; same machine different processes inconsistent. Test each upstream separately, confirm timeout is real packet loss, server busy, or TCP fallback.
for server in 10.43.0.10 10.43.0.11; do
dig @"$server" api.example.com A +time=2 +tries=1 | tail -8
done
resolvectl statistics
ss -uapn | head -30
journalctl -u systemd-resolved --since '30 minutes ago'Fix upstream stability or network loss first, then local/app cache. Single dig Query time ≠ real request percentiles; need multiple samples distinguishing cold/hot cache. Post-fix compare upstream error rate, app request error rate, tail latency over cache expiry period.
Four Boundaries of Resolution Path
glibc ndots rules don't directly apply to all runtimes: Go may use cgo or pure Go, Java has own cache, browsers/proxies may resolve independently. getent via NSS closer to most system programs but doesn't guarantee exact app simulation.
DNS query success ≠ TCP/TLS/HTTP success; troubleshooting records must separate three phases.
Host, container, Pod perspectives may use different resolv.conf and network namespaces; sampling must be in faulting process's execution environment.
Case 6: K8s Cluster Domain OK, External Domains Intermittent SERVFAIL
Internal short names success only proves Pod→DNS service partial path; external domains need recursive forward, upstream DNS, egress network. CoreDNS may serve cluster.local fine while forward plugin returns SERVFAIL due to upstream unreachable, timeout, or forwarding loop. Collect both internal Service FQDN and external domain, test A/AAAA, preserve DNS response codes and server log timeline.
kubectl -n kube-system get pods -l k8s-app=kube-dns -o wide
kubectl -n kube-system get svc kube-dns -o wide
kubectl -n kube-system get configmap coredns -o yaml
kubectl -n kube-system logs -l k8s-app=kube-dns --since=15m --max-log-requests=10 | tail -100Check CoreDNS pod restarts, CPU saturation, forward target pointing to address that re-queries cluster domain. Verify node egress to upstream DNS UDP/TCP 53; some envs block egress 53 → internal works, external fails. If upstream uses DoT/DoH, consider ports and health checks separately. Validate not single nslookup; cover faulting pod's node, another node, and direct upstream DNS as control. Change Corefile: syntax check + monitoring + rolling batches, ensure internal Service queries not collateral.
Case 7: Resolution Correct but Connects to Wrong Service
Migration: old/new instances coexist. dig returns IP matching record, but app configured proxy or /etc/hosts override → connects to old service. Or multiple A records with some stale; single dig seeing one healthy IP doesn't prove all clients hit good address. Must separate: DNS result, address selection, connection destination, server logs — not blanket "DNS pollution".
getent ahosts api.example.com
dig api.example.com A +short
sed -n '/^hosts:/p' /etc/nsswitch.conf
rg -i 'api.example.com' /etc/hosts
curl -v --connect-timeout 5 https://api.example.com/ -o /dev/null
env | rg -i '(^|_)https?_proxy=|no_proxy='Capture request trace ID, confirm on target service logs if request received. If DNS round-robin includes multiple IPs, use curl --resolve per address preserving original domain/SNI. Cache TTL, client connection reuse, proxy DNS resolution location all affect switch speed. Pre-lower TTL still can't guarantee instant client update; keep old endpoint or gateway forwarding transition window.
Case 8: Retries Amplify DNS Query Storm
App hits transient resolution failure, immediate retry. Each request triggers multiple search suffixes × A/AAAA × upstream retries → DNS QPS rises during failure. Server load appears high, client P99 high. Scaling DNS without fixing app retry backoff just exhausts new capacity. Combine app's actual resolver and cache strategy; set capped exponential backoff. Monitor request count, actual DNS query count, timeouts, SERVFAIL separately.
sudo timeout 20 tcpdump -ni any -c 300 'udp port 53 or tcp port 53'
resolvectl statisticsFocus on short-window failure rate during incident, not average latency; cached successes mask failing requests timing out at app layer. Validate by gradually reducing injected error rate, confirm app retry no longer amplifies QPS, and business end-to-end error rate drops with upstream recovery. Negative cache TTL implementation varies; don't claim changing
resolv.conf timeoutsolves all upstream issues.
Case 9: DNSSEC Validation Failure & Time Sync
DNSSEC-validating recursive resolver may return SERVFAIL due to signature chain, expiry, or system time anomaly; non-validating query succeeds. Only implicate clock if query path confirmed DNSSEC-enabled and logs show validation error. Ordinary glibc A/AAAA query doesn't fail solely due to "time out of sync".
timedatectl status
resolvectl status
dig @<DNS_SERVER_IP> example.com A +dnssecCompare same domain via known healthy recursive vs faulty recursive; check response codes and logs. Before time sync, verify NTP service status and host policy; large production time jumps may affect DB, certs, schedulers — separate control. Post-fix verify with original faulty recursive across record types, confirm app connectivity restored. dig +dnssec returning RRSIG ≠ local validation completed.
Case 10: System Resolver vs Proxy Resolver Boundary
App sets HTTPS_PROXY; external domain resolution may be done by proxy or locally, depending on proxy protocol and client impl. Pod getent hosts example.com works, curl via proxy fails resolution; fixing Pod CoreDNS may be useless. First curl -v to see if connecting to proxy, then check proxy host's resolution environment and egress.
env | rg -i 'http_proxy|https_proxy|all_proxy|no_proxy'
curl -v --connect-timeout 5 https://example.com/ -o /dev/null
curl --noproxy '*' -v --connect-timeout 5 https://example.com/ -o /dev/nullLocal proxy disable failing to connect doesn't prove DNS fault; may be egress policy mandating proxy. Record separately: client→proxy TCP connect, proxy resolves target domain, proxy→target connect, proxy return code. NO_PROXY suffix matching rules vary by client; watch case and port.
Case 11: TCP/53 Blocked, UDP Large Response Can't Fallback
Large DNS responses may be truncated in UDP, requiring client TCP/53 retry. Network policy allowing only UDP 53 → small answers work, TXT/DNSSEC/large responses fail. Test same upstream with UDP and TCP, compare TC bit and failure; check egress and CoreDNS connection logs.
dig @<DNS_SERVER_IP> example.com TXT +time=2 +tries=1
dig @<DNS_SERVER_IP> example.com TXT +tcp +time=2 +tries=1If EDNS or fragmentation loss, symptoms may relate to MTU; don't just tweak ndots on large response failure. Validate across A, AAAA, TXT, DNSSEC records and real business requests ensuring TCP fallback path open. Network policy change: only open UDP/TCP 53 to target recursive, not all node egress ports.
Pre-Publish Fact Check
Article's "DNS slow" must specify metric: single upstream Query time, system resolution latency, or HTTP time_namelookup. Three numbers involve different caches/paths; not interchangeable. Examples must use reserved addresses and example domains; no production DNS IPs, internal domains, packet captures. Final change report: anomaly start time, affected nodes/Pods, upstream health, request volume/failure rate, precise root cause, rollback criteria, re-test sample size and observation window.
Advanced: Splitting Resolution Latency into Client, CoreDNS, Upstream
When Pod getent slow, must separate: resolver library local search candidates, Pod→CoreDNS network, CoreDNS→upstream recursive. Record in Pod namespace: query name, type, response code, total latency. From node/server: CoreDNS request volume, cache hits, forwarded requests, upstream errors/timeouts. If only one Pod slow while same-node others OK → that Pod's resolv.conf, process cache, container resource pressure, local proxy. If all Pods on node slow but other nodes OK → kube-proxy/eBPF, NodeLocal DNSCache, node→DNS Service IP path.
kubectl exec -n <namespace> <pod> -- cat /etc/resolv.conf
kubectl exec -n <namespace> <pod> -- getent ahostsv4 api.example.com
kubectl -n kube-system get pods -o wide | rg 'coredns|node-local-dns'
kubectl -n kube-system get svc kube-dns -o wideSampling must cover failure window. CoreDNS healthy ≠ all upstreams healthy; forward may switch between multiple upstreams; one bad node causes long tail. If upstream latency high, test each upstream directly. If CoreDNS shows fast processing but client slow, check Service forwarding or client retry. Conclusion table must list "client observed total latency" and "server processing latency" side by side; difference points to network or queueing.
Advanced: Reading DNS Requests/Responses in Packet Captures
One logical query may generate multiple packets due to search domains, A/AAAA, retries, TCP fallback. Correlate by transaction ID, question name, source port, DNS server — don't just count port 53 packets and call it app query count. NXDOMAIN: check if request was search-expanded candidate; candidate fails then absolute name may succeed. NOERROR: verify answer not empty; no A but AAAA or CNAME is different case.
sudo timeout 30 tcpdump -ni any -vv -s 0 -c 200 'udp port 53 or tcp port 53'
dig @<DNS_SERVER_IP> api.example.com. A +time=2 +tries=1
dig @<DNS_SERVER_IP> api.example.com. AAAA +time=2 +tries=1Production captures: limit time, filter targets, save cautiously (internal domains may leak business info). If app uses DoH/DoT, this method sees only proxy/encrypted traffic; problem names from corresponding service logs. Fault report retains sanitized query name, qtype, rcode, server, latency for reproducibility.
Advanced: Negative Caching & Record Change Propagation
Record TTL is just one cache expiry in authoritative/recursive chain. NSS cache, systemd-resolved, nscd, JVM, browser, proxy, long-lived connections make cutover appearances differ. Post-record-change, a machine "still connects old IP" ≠ upstream DNS not updated: may hold cached response, or resolved new IP but reuses old TCP connection. First simultaneously record: resolution result, actual connection remote address, app connection pool and cache layers.
getent ahosts api.example.com
resolvectl query api.example.com
ss -tnp | rg ':443'
dig @<KNOWN_RESOLVER> api.example.com A +noall +answerCache clearing has business cost; mass simultaneous flush may surge upstream. Plan transition: old target overlap period, configure new/old parallel health checks. Validation: not just wait TTL seconds, observe app logs for actual connection target and error rate; long connections need business-controllable refresh strategy.
Advanced: Split-View Internal/External Same Domain
Internal DNS may resolve api.example.com to private IP; public recursive returns public IP. Manually switching resolv.conf nameserver to public DNS might "fix" one external resolution but break internal service access or route via wrong egress. Troubleshoot: first clarify which view workload should use; record same-name results from office, VPN, Pod, host — don't judge internal DNS by public dig.
getent ahostsv4 api.example.com
dig @<INTERNAL_RESOLVER> api.example.com A +short
dig @<EXTERNAL_RESOLVER> api.example.com A +shortIf split-view by design, fix internal resolver records or forward policy. Private IP unreachable from public is normal. Validation: confirm resolution to expected zone target, and actual TCP/TLS succeeds; DNS layer "has answer" is only first gate.
Minimal Evidence Package for Shift Handover
Post DNS incident retain at minimum: affected hosts/Pods/nodes, specific query domain and record type, client runtime and proxy usage, resolv.conf and NSS config, occurrence time, upstream resolver, upstream and client response codes and latencies, packet captures or logs, fix action, rollback method, post-fix multi-round validation. Monitoring: watch P95/P99, failure rate, upstream QPS — not average Query time replacing tail latency. For intermittent failures, keep low-load continuous probe with minute-level samples to see trends and anomaly windows, not rely on single manual dig.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Ops Community
A leading IT operations community where professionals share and grow together.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
