Operations 77 min read

DNS Troubleshooting Guide: From resolv.conf to Kubernetes - Resolving Intermittent Failures and Latency

A comprehensive guide to diagnosing DNS resolution latency and failures, covering resolv.conf configuration, ndots and search domain pitfalls, caching layers, Kubernetes DNS specifics, and practical troubleshooting steps with command-line tools.

Ops Community
Ops Community
Ops Community
DNS Troubleshooting Guide: From resolv.conf to Kubernetes - Resolving Intermittent Failures and Latency

Problem Background

DNS resolution is an often overlooked "hidden bomb" that can take down entire services. Common symptoms include applications occasionally hanging for tens of seconds then failing (retry succeeds), upstream DNS changes causing intermittent resolution, short names resolving to wrong hosts due to search domains, container resolution delays from ndots:5, and sudden latency spikes after upgrades.

Applicable Scenarios

Application domain resolution slow, timeout, intermittent, or occasional failures

Behavior changes after modifying resolv.conf, /etc/nsswitch.conf, /etc/hosts, or DNS configs

Internal services using short names ( db, redis) resolving to wrong hosts or high latency

Container/K8s Pod resolution behavior differs from bare metal

Need to distinguish DNS issues from network/TLS/connection issues

Need to unify DNS config across machines with validation and rollback

Prerequisite: network layer reachable (ping target IP works). If IP unreachable, fix connectivity first.

Core Knowledge

Resolution Flow

Application calls getaddrinfo() or similar

Query order determined by

/etc/nsswitch.conf
hosts:

line (commonly files dns)

DNS queries follow

/etc/resolv.conf
nameserver

entries sequentially

Query name handling depends on ndots and search (direct FQDN vs search domain suffixes)

Results may be cached by nscd, systemd-resolved, or application caches

Each nameserver has timeout (default 5s) and attempts (default 2)

Key resolv.conf Fields

nameserver 10.0.0.53          # up to 3, tried in order (or rotate with options rotate)
search example.com sub.example.com  # multiple search domains
options timeout:1 attempts:2 ndots:1 rotate
nameserver

: max 3, tried sequentially unless rotate set search / domain: append suffixes to non-FQDN names ndots: dots threshold to treat as FQDN (default 1) timeout: per-nameserver query timeout seconds (default 5) attempts: retries per nameserver (default 2) rotate: round-robin across nameservers

ndots Deep Dive

ndots:N

means names with ≥N dots are queried as FQDN directly. Default ndots:1. With search example.com: host.example.com (1 dot) → direct query host (0 dots) → tries host.example.com then host Kubernetes often sets ndots:5, causing short names and even single-dot names to attempt multiple search suffixes, each adding latency and potential timeouts. This is the #1 suspect for container resolution slowness.

search Domain Pitfalls

Search suffixes can collide with external domains (e.g., db → db.prod.example.com conflicting with public same name), causing "resolves but to wrong host". Unresolvable search domains waste timeout cycles per query.

/etc/hosts & nsswitch.conf

/etc/hosts

: static mappings, highest priority if files first in hosts: line

/etc/nsswitch.conf
hosts:

determines lookup order (e.g., files dns)

Check /etc/hosts for pollution when results mismatch expectations

Cache Layers

nscd

(legacy) / systemd-resolved (modern): cache resolution results

Application-level caches (Java, Go, Node each have own strategies)

Config changes not reflecting? Suspect cache first

Upstream & Recursion

Configured nameservers are usually recursive resolvers (VPC DNS, cloud DNS, 8.8.8.8). Slowness may be upstream or path to upstream. Use dig +trace to expand recursive path and see per-hop latency.

Overall Troubleshooting Approach

1. Confirm DNS issue (compare IP direct connect)
2. Check local resolver config (resolv.conf / nsswitch.conf / hosts)
3. Use dig to reproduce and separate local config from upstream
4. Verify ndots/search rewrite effects
5. Check timeout/retry/cache layers (timeout, attempts, nscd/resolved)
6. Verify time (NTP) and multiple nameserver order/failures
7. Fix + validate
8. Rollback plan + prevention (monitoring & cache strategy)

Practical Troubleshooting Steps

Step 1: Confirm DNS vs Connectivity

# Test connectivity directly with IP
ping -c 3 <targetIP>
curl -v --connect-timeout 3 http://<targetIP>/

# Measure resolution time
time getent hosts <domain>
time nslookup <domain>
getent hosts

slow/fails → DNS issue, continue getent hosts fast, IP curl works → issue may be in app's resolver/cache/proxy

IP unreachable → network issue, not DNS

Note: getent hosts uses glibc stack (hosts+DNS), dig queries DNS directly bypassing /etc/hosts. Different behaviors help isolate layer.

Step 2: Check Local Config

cat /etc/resolv.conf
cat /etc/nsswitch.conf | grep hosts
cat /etc/hosts

Are nameserver IPs reachable and appropriate?

Is hosts: order changed unexpectedly?

Any problematic entries in /etc/hosts?

If resolv.conf managed by DHCP/systemd-resolved, edit source not the symlink

Step 3: Use dig to Isolate Local Stack vs Upstream

# Direct query to specific server, bypass search
dig @10.0.0.53 <FQDN>

# Measure latency
dig @10.0.0.53 <FQDN> +time=2 +tries=1

# See search expansion for short name
dig <shortname> +search +time=2

# Trace recursive path
dig +trace <FQDN>
dig @server FQDN

fast + correct → upstream OK, problem local (ndots/search/order/cache) dig @server FQDN slow/timeout → upstream or path issue dig shortname +search shows actual expanded queries +trace reveals per-hop latency (root, TLD, authoritative)

Step 4: Verify ndots & search Rewrites

cat /etc/resolv.conf | grep -E 'ndots|search|options'
getent hosts <shortname> || true
dig <shortname> +search +short

With ndots:5 and search a b c, short name tries name.a, name.b, name.c, name sequentially. Each adds latency.

Optimization: lower ndots to 1 or 0 for FQDN-heavy apps, trim search domains, or use trailing dot ( name.example.com.) for absolute names.

Step 5: Timeouts, Retries & Cache

grep -E 'timeout|attempts|rotate' /etc/resolv.conf
systemctl status nscd systemd-resolved 2>/dev/null
ls -l /etc/resolv.conf
readlink -f /etc/resolv.conf
resolvectl status 2>/dev/null | head

Default timeout:5 attempts:2 can yield tens of seconds worst-case, amplified by search variants

Config change no effect? Clear cache: systemctl restart nscd / systemctl restart systemd-resolved / app restart

If systemd-resolved manages resolv.conf (symlink to /run/systemd/resolve/...), edit /etc/systemd/resolved.conf not the symlink

Step 6: NTP & Multiple Nameservers

timedatectl
timedatectl status | grep -E 'synchronized|NTP'
grep nameserver /etc/resolv.conf

Large time skew affects DNSSEC/TLS/recursive policies → sync time first

First nameserver unreachable/slow drags all queries (unless rotate). Use options rotate or put reliable server first

Step 7: Fix & Validate

Fix resolv.conf → retest with dig @server FQDN +time=2 +tries=1 and getent hosts FQDN If systemd-resolved: edit /etc/systemd/resolved.conf ( DNS=, Domains=) then systemctl restart systemd-resolved Adjust ndots /trim search for short-name scenarios

Validation must cover: getent hosts (glibc), dig (upstream), actual app requests — all fast/normal.

Common Commands Cheatsheet

# Resolution & timing
dig @server name
dig name +short
dig name +trace
getent hosts name
nslookup name
host name

# Config
cat /etc/resolv.conf
cat /etc/nsswitch.conf | grep hosts
cat /etc/hosts
ls -l /etc/resolv.conf

# Cache & services
systemctl status nscd systemd-resolved
systemctl restart nscd systemd-resolved
resolvectl status

# Time
timedatectl status

# Packet capture
tcpdump -i eth0 -n port 53 -c 50

Root Cause Deep Dives

7.1 Nameserver Unreachable / Wrong Order

cat /etc/resolv.conf | grep nameserver
for ns in $(awk '/nameserver/{print $2}' /etc/resolv.conf); do
  echo "== $ns =="; timeout 2 nc -vz "$ns" 53 2>&1 || echo "$ns:53 unreachable"
done

Fix: move reachable/fast servers first, add rotate, or remove unreachable.

7.2 High ndots Latency

cat /etc/resolv.conf | grep ndots
# or
resolvectl status

time dig shortname +search +time=2 | grep -E 'time|status'

Fix: lower ndots, trim search, or use trailing dot. Validate with time getent hosts name.example.com. vs time getent hosts name.example.com.

7.3 Search Domain Collision

getent hosts <shortname>
dig <shortname> +search

Fix: remove extra search domains, use FQDN, or pin in /etc/hosts (but record it to avoid migration issues).

7.4 Cache Staleness

systemctl restart nscd
systemctl restart systemd-resolved
getent hosts <FQDN>

7.5 Time Skew

timedatectl
timedatectl set-ntp on
# or chrony/ntpdate

7.6 Application-Layer Caches (Java, Go, Node)

Java: networkaddress.cache.ttl JVM property. Go: net.Dialer and LOOKUP env. Node: process-internal DNS cache. If system dig works but app fails, check app cache/TTL/restart.

7.7 DNSSEC Validation Failures

dig @server FQDN
# SERVFAIL (not NXDOMAIN) suggests DNSSEC issue

Verify with DNSSEC disabled or public recursive; coordinate with upstream.

Configuration Examples

8.1 Internal FQDN-Heavy Apps

nameserver 10.0.0.53
nameserver 10.0.1.53
search example.com
options timeout:1 attempts:2 ndots:1
ndots:1

lets single-dot names query directly as FQDN; short names use search.

8.2 Short-Name Heavy Environments

nameserver 10.0.0.53
search prod.example.com staging.example.com
options timeout:2 attempts:2 ndots:1 rotate
rotate

balances across nameservers.

8.3 Kubernetes/Container Notes

Pod dnsConfig may override ndots (often default ndots:5). Adjust via spec.dnsConfig.options (e.g., ndots:1), but evaluate impact on other resolutions. Example:

dnsPolicy: ClusterFirst
dnsConfig:
  options:
    - name: ndots
      value: "1"
    - name: single-request-reopen

Note: single-request-reopen behavior varies by resolver (glibc/musl/app); test before adopting.

Log & Metrics Observation

DNS often lacks dedicated logs; rely on active probes and packet capture. Most reliable: tcpdump -i any -n 'udp port 53' -c 100 and tcp port 53. Observe query destination, per-query latency, retries, failures ( [.] root marker), SERVFAIL.

If systemd-resolved: resolvectl statistics shows hit rate and latency.

Baseline: LAN resolution typically milliseconds; cross-network/public recursive tens to hundreds of ms. Alert on sustained >hundreds ms, timeouts, SERVFAIL spikes, duplicate queries.

Decision Tree (Troubleshooting Path)

Domain resolution slow/fail?
A. IP direct connect?
  No → not DNS, check network
  Yes → next
B. dig @server FQDN OK?
  No → upstream/path issue
  Yes → next
C. getent hosts FQDN OK?
  No → local config/stack/cache/time
  Yes → next
D. Short name vs FQDN same?
  Different → focus ndots/search
  Same → check cache/timeout/multi-nameserver
E. Packet capture latency/failure where?
  Local early hops → config/rewrite
  Upstream no response → upstream/firewall UDP53

Risk Warnings

Modifying /etc/resolv.conf is system config change; backup first

Manual edits may be overwritten by DHCP/systemd-resolved; use proper channels

Removing search domains may break services relying on short names

Lowering ndots may cause previously search-handled short names to fail as bare queries rotate changes query behavior; validate before production

Restarting nscd/systemd-resolved clears global cache — brief jitter, do off-peak

Pinning in /etc/hosts bypasses DNS; causes "migrated but still resolves old IP" issues — document

Always backup before delete/overwrite; provide rollback content

Production changes: backup, canary, rollback, record

Validation Checklist

Backup exists:

ls -l /etc/resolv.conf.bak*
dig @server FQDN

returns correct RCODE quickly getent hosts FQDN returns fast

App actual connections normal

Packet capture confirms query count/latency as expected (short names no extra variants)

Multiple nameserver order/rotate behaves as expected

One cycle post-change: no alerts, no new errors

Rollback Procedure

# Backup
cp /etc/resolv.conf /etc/resolv.conf.bak.$(date +%F)

# Rollback (restore pre-change file)
cp /etc/resolv.conf.bak.$(date +%F) /etc/resolv.conf

# If systemd-resolved mode: restore resolved.conf then restart
# Record original resolve config, rollback then systemctl restart systemd-resolved

If restart nscd caused jitter: wait for self-heal or restart services sequentially

Post-rollback verify getent hosts and app recovery

Production Best Practices

Change window: low traffic, notify affected teams

Canary single node then batch

Unified config via config management to prevent drift

Multi-tenant/cluster: watch interaction with app-layer caches and Pod overrides

Monitor "resolution success rate + latency"; alert on anomalies

Document: owner, upstream DNS, dependencies, rollback method

Include NTP sync in monitoring; time drift amplifies DNS issues

Prevention: DNS in Routine Ops

Baseline: record normal resolution latency/success rate for alert thresholds

Metrics: DNS query success rate, avg latency, SERVFAIL count (per actual exporter)

Quarterly audit: resolv.conf consistency, prune stale search domains, check expired /etc/hosts entries

Change log: every DNS/nameserver/ ndots change recorded

Drill: periodic "kill nameserver + rollback" exercise for critical domains

Full Case Study: Container Intermittent Timeout → ndots

Symptom: Microservice calling third-party occasionally hangs 10+ seconds, retry succeeds.

Troubleshoot: IP direct connect OK. dig @vpc-dns FQDN +time=2 fast. But getent hosts for internal short name slow. /etc/resolv.conf shows options ndots:5 search prod.example.com staging.example.com. Short name tries multiple search variants, each failed path adds wait.

Fix: For "mostly FQDN" service, lower ndots to 1 or use FQDN with trailing dot. Canary single node then batch.

Validate: time getent hosts FQDN and packet capture confirm query count drops to 1, latency back to milliseconds. Rollback: record original resolv /override config, restore on anomaly.

Root cause: ndots:5 × multiple search domains × timeout stacking. Lesson: check ndots / search rewrite impact before staring at nameservers.

dig Subcommands Deep Dive

17.1 Basic Query & RCODE

dig @10.0.0.53 example.com A

Read status ( NOERROR / NXDOMAIN / SERVFAIL), ANSWER TTL, Query time (most direct "how long").

17.2 Force No Cache, No Search

dig +norecurse example.com   # non-recursive, authoritative view
dig example.com +short       # just results
dig example.com +noall +answer
+norecurse

on recursive server returns NOERROR but no ra flag; shows if authoritative side has issue.

17.3 Trace Recursive Path

dig example.com +trace

Shows each hop from root to authoritative with Query time. Slow hop = latency source.

17.4 Query Type & TCP

dig @server example.com CNAME
dig @server +tcp example.com

UDP default; some envs filter UDP or need TCP for large responses. +tcp tests TCP path.

17.5 Timeout & Retry Control

dig @server example.com +time=2 +tries=1

Simulates "fail fast" to see if network issues cause long hangs.

17.6 Batch Queries

for n in db redis cache; do echo "== $n =="; dig $n +short; done

Quick scan which short names slow/fail.

FAQ

19.1 Why dig fast but getent hosts slow?

dig

queries specified nameserver directly, may skip search expansion and some cache layers. getent hosts uses system stack: /etc/hosts, nsswitch, ndots/search, /etc/resolv.conf. Discrepancy points to problematic layer.

19.2 How ndots affects speed?

ndots

decides dot threshold for FQDN treatment. Below threshold, search suffixes tried sequentially; each failed variant consumes attempt/wait. ndots:5 makes short names try many variants. Fix: caller uses name. (trailing dot) or ndots:1.

19.3 resolv.conf change not effective?

Likely managed by systemd-resolved or DHCP. readlink -f /etc/resolv.conf shows symlink; edit resolved config source, not symlink target.

19.4 Same domain sometimes works sometimes not?

One nameserver unreachable, timeout too small, or load spread to slow backend. Loop dig multiple times to see failure rate; add rotate or reorder nameservers.

19.5 Should I pin internal domains in /etc/hosts?

Can but cautiously. /etc/hosts highest priority, bypasses DNS. Solves "can't resolve internal" but creates "IP changed but still resolves old" risk. If used: document, limit scope, monitor for cleanup.

19.6 Does time skew affect resolution?

Yes. Large skew breaks DNSSEC validation and some recursive policies. timedatectl check; run NTP/chrony.

19.7 App cached DNS?

Java/Go/Node have own DNS caches or resolver strategies. /etc/hosts and resolv.conf changes not picked up is normal process-internal cache. Fix: adjust app TTL or restart app; not in system config.

DNS Change Compliance Template

Change Title: app-01 adjust DNS search domains and ndots
Reason: Container short-name resolution latency from ndots:5 + multiple search
Machines: app cluster (canary app-01 → two batches)
Pre-check: dig/getent baseline latency; confirm /etc/resolv.conf ownership (resolved?)
Action: backup original → apply; canary one → validate getent/dig/app → batch
Files: /etc/resolv.conf (or systemd-resolved source)
Validation: getent latency drop + dig query count down + app no errors
Rollback: cp backup restore + systemctl restart nscd/resolved; on anomaly revert
Impact: all name resolution on machine; declare affected services
Approver/Executor: ____

Monitoring Probe Script

Read-only probe /usr/local/bin/dns-health.sh collects key FQDN resolution latency and status, cron-able, monitoring-ingestible.

#!/usr/bin/env bash
# DNS health probe (read-only)
NAME=${1:-www.example.com}
SERVERS=(${@:2})
[ ${#SERVERS[@]} -eq 0 ] && SERVERS=("10.0.0.53" "119.29.29.29")

for ns in ${SERVERS[@]}; do
  out=$(dig @"$ns" "$NAME" +time=2 +tries=1 2>/dev/null)
  st=$(printf '%s
' "$out" | awk '/status:/{print $4}')
  qt=$(printf '%s
' "$out" | awk '/Query time/{print $4}')
  echo "$(date +%F\ %T) $ns $NAME status=$st qtime=${qt}ms"
done

# System stack perspective
start=$(date +%s%3N)
getent hosts "$NAME" >/dev/null 2>&1; rc=$?
end=$(date +%s%3N)
echo "system-stack name=$NAME rc=$rc ms=$((end-start))"

Usage:

sudo bash /usr/local/bin/dns-health.sh www.example.com 10.0.0.53

. Non-intrusive; log to cron or push metrics. Alert on non- NOERROR status or qtime exceeding baseline.

Rollback Nuances: Preventing Recurrence

Manual /etc/resolv.conf edit later overwritten by DHCP/resolved → rollback must align with source, else re-overwritten

systemd-resolved rollback = restore original resolved.conf then restart

App-layer TTL changes need rollback to original config /etc/hosts pinned records rollback = delete lines

Rollback isn't just "revert values in 3 seconds"; must ensure source won't re-apply. Often overlooked, causes re-recurrence.

Minimal Monitoring & Alerting Loop

Collect: scheduled dns-health.sh → push qtime / status / rc to logs/metrics

Metrics: resolution success rate, avg latency, SERVFAIL / NXDOMAIN counts (per collector)

Thresholds: sustained 0% success or latency > baseline N× → alert; tune per business cycle, no absolute hard numbers

Linkage: alert triggers "dig → getent → tcpdump" mainline auto/manual reproduction

IPv6 & DNS: Don't Only Debug IPv4

Many "slow/weird" resolutions involve IPv6 interference. Check app/system v6 resolution path:

ip -6 addr
cat /etc/gai.conf | grep -v '^#' | head

dig @10.0.0.53 www.example.com AAAA +time=2
getent ahosts www.example.com | head

Common pitfalls: getent ahosts returns mixed A/AAAA; app may stall waiting for non-existent AAAA ("AAAA stall") /etc/gai.conf address selection (RFC 6724) reorders IPv4 vs IPv6 preference

IPv6 firewall ( ip6tables /nft ip6) unconfigured → v6 connections blocked/loopback

Logic: App prefers AAAA but no v6 connectivity → extra wait. Adjust gai.conf order or force v4 in app. Validate v4/v6 separately with ping6 / curl -6/-4. Backup gai.conf before changes; canary validate.

DoH/DoT & Forwarders Reminder

Modern stacks use dnscrypt / DoH / DoT, sslh, unbound, dnsmasq. If resolv.conf nameserver is 127.0.0.53 / 127.0.0.1, real upstream is in forwarder config.

grep nameserver /etc/resolv.conf
cat /etc/systemd/resolved.conf 2>/dev/null
cat /etc/unbound/unbound.conf 2>/dev/null | grep -iE 'forward|tcp|tls'

If nameserver is localhost stub, editing resolv.conf nameserver may be ineffective; edit forwarder component. Explains "I changed resolv.conf but nothing changed".

Quantitative: qtime, Hit Count, Failure Rate Together

Single slow query ≠ incident. Quantify. Collection pseudocode:

for i in $(seq 1 20); do
  dig @10.0.0.53 www.example.com +time=1 +tries=1 2>/dev/null
done | grep -c 'NOERROR'  # success count

Logic: Success rate significantly <100% or latency exceeds baseline → formal investigation. Transient single slow query not incident; look at trends and multiple metrics (qtime, SERVFAIL count, reachability) jointly.

Cross-Machine DNS Config Consistency

Many machines → DNS drift causes "this works, that doesn't". Use config management:

- name: Deploy resolv.conf
  ansible.builtin.copy:
    content: |
      nameserver 10.0.0.53
      search example.com
      options timeout:2 attempts:2 ndots:1
    dest: /etc/resolv.conf
  notify: restart resolv-manager

If all resolv.conf consistent but differences persist → network/upstream/regional diff, don't blame DNS alone. Rollback = restore old resolv.conf and config.

One-Page DNS Troubleshooting Checklist

☐ 1. IP direct connect OK? (ping/curl target IP)
☐ 2. dig @server FQDN fast and NOERROR?
☐ 3. getent hosts FQDN fast?
☐ 4. resolv.conf nameservers reachable and first?
☐ 5. systemd-resolved/DHCP managed? (readlink)
☐ 6. ndots too high? search domains excessive?
☐ 7. Short name expanded to wrong domain?
☐ 8. After config change, nscd/resolved cache cleared?
☐ 9. App-layer independent cache/TTL?
☐ 10. Time sync (NTP) OK?
☐ 11. IPv6 interference?
☐ 12. Forwarder layer (unbound/dnsmasq) intercepting?
☐ 13. Packet capture (tcpdump port 53) confirms query path
☐ 14. Baseline/rollback ready

If checklist exhausted and root cause not found, issue likely further upstream in recursive/authoritative chain or requires coordination. This checklist maps 1:1 with article sections — condensed execution version.

Field Case Studies

Case 1: Resolved IP but Connection Hangs at Trying

Container curl -v https://www.example.com/ shows "Host resolved" then "Trying 203.0.113.20:443..." timeout. DNS already succeeded. Need to trace SYN egress: container → pod NIC → node → firewall/NAT → SYN-ACK.

getent ahostsv4 www.example.com
curl -4v --connect-timeout 5 https://www.example.com/
ip route get 203.0.113.20
sudo tcpdump -ni any 'host 203.0.113.20 and tcp port 443'

Only SYN retransmits → check NetworkPolicy, node egress, SNAT, upstream ACL, target service. TLS stall after handshake → verify SNI, cert, proxy, middlebox. Use curl --resolve domain:443:IP https://domain/ to preserve Host/SNI. Measure DNS/TCP/TLS phases with

curl -w '%{time_namelookup} %{time_connect} %{time_appconnect}
'

.

Case 2: Pod External Domain Latency Amplified by search Domains

Pod resolv.conf has three K8s search suffixes and ndots:5. App heavily accesses api.example.com. Under glibc, dot count <5 → tries search variants first. If each candidate times out, external domain first-resolution latency amplifies; if quick NXDOMAIN, impact small. Must measure actual queries and latency.

cat /etc/resolv.conf
getent ahosts api.example.com
dig api.example.com. A +time=2 +tries=1
sudo tcpdump -ni any -vv 'port 53'

Validate: adjust dnsConfig.options.ndots or use trailing dot for specific FQDN, compare real app (glibc/Go/Java differ) P95 resolution latency, DNS QPS, cluster short service calls. Don't blindly set ndots:0 — breaks service.namespace internal short names. Modify via workload declaration, canary, not hand-edit container file.

Case 3: dig OK but App Reports Cannot Resolve

dig @10.43.0.10 api.example.com

returns A record, but Python/Java says domain not exist. dig direct query ≠ NSS, /etc/hosts, systemd-resolved, local cache, nor proves app's network namespace matches test shell.

getent hosts api.example.com
sed -n '/^hosts:/p' /etc/nsswitch.conf
cat /etc/hosts
readlink -f /etc/resolv.conf
sudo nsenter -t <pid> -n getent hosts api.example.com
resolvectl query api.example.com

Divergence: getent fails, dig works → check NSS order, hosts override, local stub, search domains. System call OK but app fails → app DNS cache, proxy config, runtime independent resolver, container file view. Record current resolution and cache TTL before clearing; mere app restart may mask upstream issue.

Case 4: A Record OK, IPv6 Path Times Out

Domain has A and AAAA; app connects via IPv6 slowly then falls back to IPv4. DNS not failing; issue in address selection and IPv6 routing. Only dig A misses this branch.

dig api.example.com A +short
dig api.example.com AAAA +short
getent ahosts api.example.com
curl -4v --connect-timeout 3 https://api.example.com/
curl -6v --connect-timeout 3 https://api.example.com/

Cross-reference ip -6 route, container IPv6 support, egress firewall, proxy AAAA handling, target service listeners. Force IPv4 only as temporary diagnostic/recorded fallback, not delete public AAAA before root cause clear. Retest must cover IPv4/IPv6 success rates and client selection path.

Case 5: Multiple Nameservers + Cache → "Fix Not Effective"

First nameserver intermittent packet loss, second healthy; app shows periodic long tail. Resolvers differ in retry/rotate/cache strategies; same machine different processes inconsistent. Test each upstream separately, confirm timeout is real packet loss, server busy, or TCP fallback.

for server in 10.43.0.10 10.43.0.11; do
  dig @"$server" api.example.com A +time=2 +tries=1 | tail -8
done
resolvectl statistics
ss -uapn | head -30
journalctl -u systemd-resolved --since '30 minutes ago'

Fix upstream stability or network loss first, then local/app cache. Single dig Query time ≠ real request percentiles; need multiple samples distinguishing cold/hot cache. Post-fix compare upstream error rate, app request error rate, tail latency over cache expiry period.

Four Boundaries of Resolution Path

glibc ndots rules don't directly apply to all runtimes: Go may use cgo or pure Go, Java has own cache, browsers/proxies may resolve independently. getent via NSS closer to most system programs but doesn't guarantee exact app simulation.

DNS query success ≠ TCP/TLS/HTTP success; troubleshooting records must separate three phases.

Host, container, Pod perspectives may use different resolv.conf and network namespaces; sampling must be in faulting process's execution environment.

Case 6: K8s Cluster Domain OK, External Domains Intermittent SERVFAIL

Internal short names success only proves Pod→DNS service partial path; external domains need recursive forward, upstream DNS, egress network. CoreDNS may serve cluster.local fine while forward plugin returns SERVFAIL due to upstream unreachable, timeout, or forwarding loop. Collect both internal Service FQDN and external domain, test A/AAAA, preserve DNS response codes and server log timeline.

kubectl -n kube-system get pods -l k8s-app=kube-dns -o wide
kubectl -n kube-system get svc kube-dns -o wide
kubectl -n kube-system get configmap coredns -o yaml
kubectl -n kube-system logs -l k8s-app=kube-dns --since=15m --max-log-requests=10 | tail -100

Check CoreDNS pod restarts, CPU saturation, forward target pointing to address that re-queries cluster domain. Verify node egress to upstream DNS UDP/TCP 53; some envs block egress 53 → internal works, external fails. If upstream uses DoT/DoH, consider ports and health checks separately. Validate not single nslookup; cover faulting pod's node, another node, and direct upstream DNS as control. Change Corefile: syntax check + monitoring + rolling batches, ensure internal Service queries not collateral.

Case 7: Resolution Correct but Connects to Wrong Service

Migration: old/new instances coexist. dig returns IP matching record, but app configured proxy or /etc/hosts override → connects to old service. Or multiple A records with some stale; single dig seeing one healthy IP doesn't prove all clients hit good address. Must separate: DNS result, address selection, connection destination, server logs — not blanket "DNS pollution".

getent ahosts api.example.com
dig api.example.com A +short
sed -n '/^hosts:/p' /etc/nsswitch.conf
rg -i 'api.example.com' /etc/hosts
curl -v --connect-timeout 5 https://api.example.com/ -o /dev/null
env | rg -i '(^|_)https?_proxy=|no_proxy='

Capture request trace ID, confirm on target service logs if request received. If DNS round-robin includes multiple IPs, use curl --resolve per address preserving original domain/SNI. Cache TTL, client connection reuse, proxy DNS resolution location all affect switch speed. Pre-lower TTL still can't guarantee instant client update; keep old endpoint or gateway forwarding transition window.

Case 8: Retries Amplify DNS Query Storm

App hits transient resolution failure, immediate retry. Each request triggers multiple search suffixes × A/AAAA × upstream retries → DNS QPS rises during failure. Server load appears high, client P99 high. Scaling DNS without fixing app retry backoff just exhausts new capacity. Combine app's actual resolver and cache strategy; set capped exponential backoff. Monitor request count, actual DNS query count, timeouts, SERVFAIL separately.

sudo timeout 20 tcpdump -ni any -c 300 'udp port 53 or tcp port 53'
resolvectl statistics

Focus on short-window failure rate during incident, not average latency; cached successes mask failing requests timing out at app layer. Validate by gradually reducing injected error rate, confirm app retry no longer amplifies QPS, and business end-to-end error rate drops with upstream recovery. Negative cache TTL implementation varies; don't claim changing

resolv.conf
timeout

solves all upstream issues.

Case 9: DNSSEC Validation Failure & Time Sync

DNSSEC-validating recursive resolver may return SERVFAIL due to signature chain, expiry, or system time anomaly; non-validating query succeeds. Only implicate clock if query path confirmed DNSSEC-enabled and logs show validation error. Ordinary glibc A/AAAA query doesn't fail solely due to "time out of sync".

timedatectl status
resolvectl status
dig @<DNS_SERVER_IP> example.com A +dnssec

Compare same domain via known healthy recursive vs faulty recursive; check response codes and logs. Before time sync, verify NTP service status and host policy; large production time jumps may affect DB, certs, schedulers — separate control. Post-fix verify with original faulty recursive across record types, confirm app connectivity restored. dig +dnssec returning RRSIG ≠ local validation completed.

Case 10: System Resolver vs Proxy Resolver Boundary

App sets HTTPS_PROXY; external domain resolution may be done by proxy or locally, depending on proxy protocol and client impl. Pod getent hosts example.com works, curl via proxy fails resolution; fixing Pod CoreDNS may be useless. First curl -v to see if connecting to proxy, then check proxy host's resolution environment and egress.

env | rg -i 'http_proxy|https_proxy|all_proxy|no_proxy'
curl -v --connect-timeout 5 https://example.com/ -o /dev/null
curl --noproxy '*' -v --connect-timeout 5 https://example.com/ -o /dev/null

Local proxy disable failing to connect doesn't prove DNS fault; may be egress policy mandating proxy. Record separately: client→proxy TCP connect, proxy resolves target domain, proxy→target connect, proxy return code. NO_PROXY suffix matching rules vary by client; watch case and port.

Case 11: TCP/53 Blocked, UDP Large Response Can't Fallback

Large DNS responses may be truncated in UDP, requiring client TCP/53 retry. Network policy allowing only UDP 53 → small answers work, TXT/DNSSEC/large responses fail. Test same upstream with UDP and TCP, compare TC bit and failure; check egress and CoreDNS connection logs.

dig @<DNS_SERVER_IP> example.com TXT +time=2 +tries=1
dig @<DNS_SERVER_IP> example.com TXT +tcp +time=2 +tries=1

If EDNS or fragmentation loss, symptoms may relate to MTU; don't just tweak ndots on large response failure. Validate across A, AAAA, TXT, DNSSEC records and real business requests ensuring TCP fallback path open. Network policy change: only open UDP/TCP 53 to target recursive, not all node egress ports.

Pre-Publish Fact Check

Article's "DNS slow" must specify metric: single upstream Query time, system resolution latency, or HTTP time_namelookup. Three numbers involve different caches/paths; not interchangeable. Examples must use reserved addresses and example domains; no production DNS IPs, internal domains, packet captures. Final change report: anomaly start time, affected nodes/Pods, upstream health, request volume/failure rate, precise root cause, rollback criteria, re-test sample size and observation window.

Advanced: Splitting Resolution Latency into Client, CoreDNS, Upstream

When Pod getent slow, must separate: resolver library local search candidates, Pod→CoreDNS network, CoreDNS→upstream recursive. Record in Pod namespace: query name, type, response code, total latency. From node/server: CoreDNS request volume, cache hits, forwarded requests, upstream errors/timeouts. If only one Pod slow while same-node others OK → that Pod's resolv.conf, process cache, container resource pressure, local proxy. If all Pods on node slow but other nodes OK → kube-proxy/eBPF, NodeLocal DNSCache, node→DNS Service IP path.

kubectl exec -n <namespace> <pod> -- cat /etc/resolv.conf
kubectl exec -n <namespace> <pod> -- getent ahostsv4 api.example.com
kubectl -n kube-system get pods -o wide | rg 'coredns|node-local-dns'
kubectl -n kube-system get svc kube-dns -o wide

Sampling must cover failure window. CoreDNS healthy ≠ all upstreams healthy; forward may switch between multiple upstreams; one bad node causes long tail. If upstream latency high, test each upstream directly. If CoreDNS shows fast processing but client slow, check Service forwarding or client retry. Conclusion table must list "client observed total latency" and "server processing latency" side by side; difference points to network or queueing.

Advanced: Reading DNS Requests/Responses in Packet Captures

One logical query may generate multiple packets due to search domains, A/AAAA, retries, TCP fallback. Correlate by transaction ID, question name, source port, DNS server — don't just count port 53 packets and call it app query count. NXDOMAIN: check if request was search-expanded candidate; candidate fails then absolute name may succeed. NOERROR: verify answer not empty; no A but AAAA or CNAME is different case.

sudo timeout 30 tcpdump -ni any -vv -s 0 -c 200 'udp port 53 or tcp port 53'
dig @<DNS_SERVER_IP> api.example.com. A +time=2 +tries=1
dig @<DNS_SERVER_IP> api.example.com. AAAA +time=2 +tries=1

Production captures: limit time, filter targets, save cautiously (internal domains may leak business info). If app uses DoH/DoT, this method sees only proxy/encrypted traffic; problem names from corresponding service logs. Fault report retains sanitized query name, qtype, rcode, server, latency for reproducibility.

Advanced: Negative Caching & Record Change Propagation

Record TTL is just one cache expiry in authoritative/recursive chain. NSS cache, systemd-resolved, nscd, JVM, browser, proxy, long-lived connections make cutover appearances differ. Post-record-change, a machine "still connects old IP" ≠ upstream DNS not updated: may hold cached response, or resolved new IP but reuses old TCP connection. First simultaneously record: resolution result, actual connection remote address, app connection pool and cache layers.

getent ahosts api.example.com
resolvectl query api.example.com
ss -tnp | rg ':443'
dig @<KNOWN_RESOLVER> api.example.com A +noall +answer

Cache clearing has business cost; mass simultaneous flush may surge upstream. Plan transition: old target overlap period, configure new/old parallel health checks. Validation: not just wait TTL seconds, observe app logs for actual connection target and error rate; long connections need business-controllable refresh strategy.

Advanced: Split-View Internal/External Same Domain

Internal DNS may resolve api.example.com to private IP; public recursive returns public IP. Manually switching resolv.conf nameserver to public DNS might "fix" one external resolution but break internal service access or route via wrong egress. Troubleshoot: first clarify which view workload should use; record same-name results from office, VPN, Pod, host — don't judge internal DNS by public dig.

getent ahostsv4 api.example.com
dig @<INTERNAL_RESOLVER> api.example.com A +short
dig @<EXTERNAL_RESOLVER> api.example.com A +short

If split-view by design, fix internal resolver records or forward policy. Private IP unreachable from public is normal. Validation: confirm resolution to expected zone target, and actual TCP/TLS succeeds; DNS layer "has answer" is only first gate.

Minimal Evidence Package for Shift Handover

Post DNS incident retain at minimum: affected hosts/Pods/nodes, specific query domain and record type, client runtime and proxy usage, resolv.conf and NSS config, occurrence time, upstream resolver, upstream and client response codes and latencies, packet captures or logs, fix action, rollback method, post-fix multi-round validation. Monitoring: watch P95/P99, failure rate, upstream QPS — not average Query time replacing tail latency. For intermittent failures, keep low-load continuous probe with minute-level samples to see trends and anomaly windows, not rely on single manual dig.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

CoreDNStcpdumpdigresolv.confndotssystemd-resolvedDNS troubleshootinggetentKubernetes DNSsearch domains
Ops Community
Written by

Ops Community

A leading IT operations community where professionals share and grow together.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.