Kubernetes Service Access Troubleshooting: From DNS to Ingress Layer-by-Layer
A comprehensive guide to diagnosing Kubernetes service connectivity issues by systematically verifying each network layer — DNS, Ingress, Service, Endpoints, kube-proxy, Pod readiness, and NetworkPolicy — with executable commands and a real-world case study.
Problem Background
Kubernetes network paths are far longer than single-machine eras: a client request traverses load balancer, Ingress Controller, Service ClusterIP, kube-proxy iptables/IPVS rules, Pod network namespace, and finally the application process. Any layer failure manifests as "cannot connect" but root causes and fixes differ entirely. This article breaks down the chain, providing layer-by-layer verification commands, judgment criteria, and remediation actions to converge issues within ten minutes instead of blindly restarting Pods.
Typical Trigger Scenarios
A table maps change types to common misoperations and symptoms:
Application release: changed container port but not Service targetPort → 503/Connection refused
Application release: changed Pod label but not Service selector → empty Endpoints, 503
Application release: wrong readiness probe path → Pod stays NotReady, empty Endpoints
Config change: app listen address changed from 0.0.0.0 to 127.0.0.1 → Endpoints normal but unreachable
Ingress change: host typo or missing www prefix → 404 or default backend
Certificate change: TLS Secret missing or domain mismatch → browser cert error or Ingress fails to start
NetworkPolicy: new policy blocks Ingress Controller → timeout, no logs
Cluster ops: node NotReady, kube-proxy abnormal, conntrack full → intermittent or widespread failures
DNS: CoreDNS Pods all abnormal, ndots misconfiguration → cluster-internal resolution timeout
Core Knowledge Points
3.1 Complete Request Path
Client
① DNS resolve example.com
▼
Entry LB (cloud SLB / hardware LB / NodePort node)
② Forward to Ingress Controller Pod or NodePort
▼
Ingress Controller (e.g., ingress-nginx)
③ Match host+path → backend service name & port
▼
Service ClusterIP
④ kube-proxy iptables/IPVS DNAT rules
▼
Pod IP:targetPort
⑤ Enter Pod network namespace → app listening port
▼
Application ProcessEach layer's input/output can be verified independently with specific commands (dig/nslookup for DNS, curl -v for entry, kubectl describe ingress for Ingress, kubectl get svc -o yaml for Service, kubectl get endpoints for Endpoints, kubectl get pod + ss for Pod, nslookup/dig for cluster-internal DNS).
3.2 Service Types
ClusterIP (cluster-internal, service-to-service, Ingress backend), NodePort (external via node IP:port, test env), LoadBalancer (cloud LB auto-created, production ingress), ExternalName (maps to external DNS, access external services). Key: ClusterIP is virtual, not on any NIC; ping ClusterIP fails normally — use curl or nc to test port connectivity.
3.3 Endpoints & EndpointSlice
Service is just rule collection; Endpoints (legacy) and EndpointSlice (GA since 1.21, discovery.k8s.io/v1) determine actual traffic destinations. Both should be consistent; inconsistency indicates control-plane sync issue (check kube-controller-manager). Empty Endpoints means Service selector matches no Ready Pods — the most common cause of 503 errors.
3.4 Ready vs Available
A Pod enters Service backend list only if: Running (not Pending/CrashLoopBackOff), all containers pass readinessProbe, not terminating (deletionTimestamp empty). Only livenessProbe without readinessProbe means Pod considered Ready as soon as process alive, causing 502 during rollout before app initializes.
3.5 kube-proxy Modes
iptables: rule chains KUBE-SERVICES/KUBE-SVC-xxx/KUBE-SEP-xxx, linear match degrades with rule count. ipvs: kernel hash table, better performance; view with ipvsadm -Ln. Check actual mode via curl -s http://127.0.0.1:10249/proxyMode (ConfigMap may be overridden by command-line args).
3.6 CoreDNS Resolution
Pod /etc/resolv.conf typically: search default.svc.cluster.local svc.cluster.local cluster.local; nameserver 10.96.0.10; options ndots:5. ndots:5 causes queries with <5 dots (e.g., mysql, api.default) to try search domains first, generating extra NXDOMAIN queries. Fix: append trailing dot (api.example.com.) or set dnsConfig.options.ndots in Pod spec.
3.7 Ingress & IngressClass
Ingress is declarative; Ingress Controller does the work. spec.ingressClassName must match existing IngressClass or Controller ignores rule (404). Legacy kubernetes.io/ingress.class annotation coexists but behavior varies if both present and differ. Missing default backend returns Controller's built-in 404 page — distinct from app 404.
3.8 HTTP Status Code Semantics
404 → Ingress (host/path mismatch, ingressClassName error, no matching rule). 502 → Backend (Pod terminated, app crash, protocol mismatch HTTP→HTTPS). 503 → Service/Endpoints (empty Endpoints, all Pods NotReady). 504 → Backend/Network (backend timeout, NetworkPolicy drop, MTU issue). Connection refused → Transport layer (no listener, iptables REJECT). Connection timeout → Network layer (packet loss: NetworkPolicy, security group, routing).
3.9 Namespace & FQDN
Cross-namespace access requires FQDN: <service>.<namespace>.svc.cluster.local. Short name only resolves within same namespace — common pitfall when hardcoding service names in configs.
Overall Troubleshooting Approach
4.1 Layered Localization Principle
Core principle: from entry inward, verify layer by layer, never skip layers. Each layer is prerequisite for next; if Endpoints empty, kube-proxy correctness irrelevant; if targetPort wrong, checking Pod listening useless.
4.2 Golden Five Questions
Who accesses whom: client in/out cluster? domain or IP? which domain?
What symptom: timeout, refuse, status code? response body features?
When started: corresponding change time (release, scale, cert update, NetworkPolicy)?
Scope: all requests or partial? specific node/Pod only?
Previously normal? Never worked = config issue; suddenly broken = change or resource issue.
Question 5 especially critical: "never worked" usually selector/port/policy misconfig; "suddenly broken" prioritize recent changes and cluster events.
4.3 Decision Tree
Access failed
│
├─ DNS issue?
│ ├─ Yes → Check DNS / Entry LB / hosts
│ └─ No ↓
├─ Ingress returns 404?
│ ├─ Yes → Check host/path/ingressClassName/Controller logs
│ └─ No ↓
├─ Returns 503?
│ ├─ Yes → Check Endpoints/EndpointSlice empty
│ │ ├─ Empty → Check Service selector, Pod label, readinessProbe
│ │ └─ Not empty → Check targetPort vs container port
│ └─ No ↓
├─ Timeout (504/timeout)?
│ ├─ Yes → Check NetworkPolicy, security group, conntrack, MTU, node status
│ └─ No ↓
├─ Connection refused?
│ ├─ Yes → Check Pod app listening, listen address 0.0.0.0
│ └─ No ↓
└─ Cluster-internal service call failed?
├─ Yes → Check CoreDNS, ndots, NetworkPolicy, cross-namespace FQDN
└─ No → Back to entry layer, check cloud LB/NodePort chain4.4 Must-Collect Info First
Read-only commands to gather Service, Endpoints, EndpointSlice, Pod status, Ingress, IngressClass, recent events — safe for production.
Practical Steps (5.0–5.11)
5.0 Temp Diagnostic Pod
Launch netshoot Pod (nicolaka/netshoot:latest) with dig, nslookup, curl, nc, tcpdump, mtr, ip tools. Delete after use.
5.1 Confirm Fault Boundary
dig +short example.com A; nslookup example.com 8.8.8.8; identify resolved IP (cloud SLB VIP?); curl -v to see TLS handshake and response headers. Table maps anomalies (NXDOMAIN, internal IP for public client, SSL cert error, Connection refused, timeout, 404 simple page, 503 with nginx header) to layer and next step. Key: if curl returns Controller's standard error page, request reached Controller → internal issue; no connection → pre-entry issue.
5.2 Ingress Rule Match
kubectl get ingress -o wide (check ADDRESS assigned); kubectl get ingress web-api -o yaml; extract rules with jsonpath. Check: ADDRESS empty (Controller not ready or IngressClass mismatch), host mismatch (typo, missing www), ingressClassName missing/non-existent (Controller ignores rule), backend service name non-existent. Verify IngressClass with kubectl get ingressclass. Check Controller logs for request matching. Real example: ingressClassName typo (nginx-ingress vs actual nginx) caused 404 with no error logs.
5.3 Service Config Correctness
kubectl get svc -o yaml; verify selector matches Pod labels exactly (app:web-api vs app:webapi), targetPort equals container listening port (port is Service port, targetPort forwards to Pod), targetPort numeric or named (if named, must match container ports[].name). Common errors shown in YAML snippets.
5.4 Endpoints/EndpointSlice Empty?
kubectl get endpoints/endpointslice -o wide; describe endpoints; jsonpath for ready/terminating conditions. Expected: multiple ready=true, terminating=false. Anomalies: <none> (selector mismatch or no Ready Pods), partial IPs (some Pods NotReady), wrong port (targetPort error), ready=false (probe failing), Endpoints vs EndpointSlice mismatch (control-plane sync issue). Commands to cross-check labels and Pod Ready status; describe pod events for Readiness probe failed.
5.5 Why Pod Not Ready
Check Pod status, probe config, container state/restarts, events, logs (including --previous). Typical probe pitfalls: initialDelaySeconds too short; better use startupProbe (delays liveness/readiness until success) with failureThreshold/periodSeconds allowing up to 300s startup. Manually verify probe inside Pod with curl/wget. If container-internal probe works but kubelet's fails, check probe host (default Pod IP) and NetworkPolicy blocking kubelet probes.
5.6 Container App Listening?
Exec into Pod: ss -lntp / netstat -lntp. Local Address must be 0.0.0.0:port or :::port (IPv6). 127.0.0.1:port cannot receive Service DNAT traffic — common root cause. Verify Pod IP connectivity from netshoot Pod with nc -zv.
5.7 Service→Pod Forwarding Path
From netshoot Pod: curl ClusterIP:port or nc. Decision matrix: Pod IP reachable + ClusterIP reachable → Service layer OK, move up to Ingress; Pod IP reachable + ClusterIP unreachable → kube-proxy rule issue (step 5.8); both unreachable → app or Pod network issue. Check iptables rules (KUBE-SERVICES, KUBE-SVC-xxx, KUBE-SEP-xxx DNAT) or ipvsadm -Ln/--stats. Monitor kube-proxy sync metrics (sync_proxy_rules_duration_seconds, sync_proxy_rules_last_timestamp_seconds).
5.8 Cluster DNS Troubleshooting
Check CoreDNS Pods, kube-dns Service, CoreDNS ConfigMap. From test Pod: cat /etc/resolv.conf; nslookup short and FQDN; dig @10.96.0.10. Anomalies: timeout (CoreDNS down or traffic blocked), NXDOMAIN (wrong name or missing FQDN), intermittent timeout (CoreDNS under-provisioned), slow success (ndots), all Pods fail (kube-dns Endpoints empty). Optimize with dnsConfig ndots=2, timeout=2, attempts=2.
5.9 NetworkPolicy Blocking?
NetworkPolicy is whitelist: once a Pod matched by any ingress rule, it goes from default-allow to only-allow-listed. Common mistake: allow app=caller but forget Ingress Controller namespace, monitoring, etc. Correct example includes namespaceSelector for ingress-nginx and podSelector for prometheus. Verify by temporarily deleting policy (with rollback plan) or tcpdump on node (SYN without SYN-ACK = dropped).
5.10 Kernel Layer & Conntrack
Check nf_conntrack_count vs max; dmesg for "nf_conntrack: table full, dropping packet" — definitive evidence of drops causing random timeouts. Increase via sysctl (persist in /etc/sysctl.d/99-k8s-network.conf) but note kernel memory cost (~300MB per 1M entries). Check TIME_WAIT count (ss -s), port range (sysctl net.ipv4.ip_local_port_range), connection distribution per destination IP (identify client leaks). MTU mismatch: small requests work, large hang; test with ping -M do -s 1472.
5.11 NodePort & Cloud LB
Check NodePort allocation, externalTrafficPolicy (Cluster: forwards to any node's backend Pods, loses source IP; Local: only local Pods, preserves source IP but nodes without Pods fail — expected behavior). Verify per-node with curl loop. Cloud LB: confirm backend nodes, health check port/path, security group allowing health checks (failure removes node from LB → direct node works, LB fails).
Command Cheat Sheets (6.1–6.7)
Organized by phase: quick health check, Service/Endpoints, Pod/container, Ingress/Controller, DNS, Node/kernel, temp diagnostic Pod. All read-only, production-safe.
Configuration Examples (7.1–7.3)
7.1 Complete Deployment+Service+Ingress
YAML with matching labels, named containerPort (http), targetPort by name, startupProbe+readinessProbe+livenessProbe, resources, Ingress with ingressClassName, TLS secret, pathType=Prefix. Key checks: pathType required, targetPort by name stable, ingressClassName must exist.
7.2 Common Misconfigs & Symptoms
Selector case/space mismatch → empty Endpoints/503; Pod listen 127.0.0.1 → Endpoints normal but connection refused; probe path wrong → Pod NotReady/empty Endpoints; Ingress backend port != Service port → 404/503; TLS Secret wrong namespace → secret not found/HTTPS fails.
7.3 Correct NetworkPolicy
Allow Ingress Controller via namespaceSelector (kubernetes.io/metadata.name: ingress-nginx, requires 1.21+), app callers via podSelector, monitoring via podSelector. Must confirm no other callers (cron, CI) omitted.
Log & Metrics Observation (8.1–8.5)
8.1 Ingress Controller Logs
nginx access log fields: status code (503=upstream unavailable), upstream name (confirm Service), upstream addresses (empty = no Endpoints), tried backends (empty = no backends), upstream response time (0 = no connection). Empty upstream addresses → jump to Endpoints troubleshooting.
8.2 kubelet & Pod Events
Readiness probe failed, BackOff restarting, Failed to pull image, Unhealthy, FailedScheduling — each with handling direction.
8.3 Key Metrics
kube_endpoint_address_available (0 = fault), kube_endpoint_address_not_ready (>0 = unhealthy Pods), nginx_ingress_controller_requests (5xx spike), nginx_ingress_controller_request_duration_seconds (p99 spike), kubeproxy_sync_proxy_rules_duration_seconds (rising = rule pressure), coredns_dns_requests_total (spike = ndots issue), coredns_dns_responses_total by rcode (SERVFAIL/NXDOMAIN rise), coredns_dns_request_duration_seconds (p99 rise = slow resolution), container_network_receive_packets_dropped_total (>0 = network issue), node_nf_conntrack_entries vs limit (near limit = bottleneck). Alert thresholds must match business baseline.
8.4 PromQL Verification
Queries for 5xx rate, available endpoints, CoreDNS SERVFAIL rate, conntrack utilization, container network drop rate.
8.5 Log Correlation
Enable CoreDNS log plugin temporarily for DNS debugging (high volume warning).
Symptom-Based Lookup (9.1–9.7)
Concise command sequences for: 404, 503, 502, timeout (504/no response), Connection refused, cluster-internal call failure, partial node failure.
Risk Warnings (10)
High risk: delete/recreate objects (Deployment, Pods, Service) — confirm traffic, rollout history, ClusterIP change impact; prefer rollout restart. High risk: restart kube-proxy/CoreDNS — impacts whole cluster; delete sequentially with wait. Medium risk: kernel params (conntrack_max) — check memory, rollback plan, maintenance window. Medium risk: NetworkPolicy changes — backup first, delete rules causes outage. Low risk: temp diagnostic Pod — delete by exact name, not --all. tcpdump: limit packets, duration, output file, delete pcap (may contain sensitive data).
Verification After Fix (11)
11.1 Layered Verification Checklist
Commands to verify each layer: Service selector, Endpoints non-empty, all Pods Ready, Pod IP direct connect, ClusterIP connect, Ingress with --resolve (bypass DNS).
11.2 Per-Backend Pod Verification
Loop over all Pod IPs with nc; expect all OK. FAIL on any means ~1/N request failures — typical partial-failure pattern.
11.3 Continuous Observation
Watch Controller 5xx for 10-30 min; small-load curl loop (50 requests) expect all 200.
11.4 Business-Side Verification
Real user flow: login, query, order — often reveals issues technical metrics miss.
Rollback Plans (12)
Deployment: rollout undo (only Pod template; Service/Ingress/ConfigMap separate). Config: backup YAML, strip runtime fields (resourceVersion, uid, creationTimestamp) before apply. NetworkPolicy: backup, delete new policy or --all (emergency only, restores default-allow). CoreDNS: backup ConfigMap, apply backup, rollout restart deployment coredns. Decision boundary: immediate rollback if 5xx exceeds threshold or core transaction affected; can debug if non-core with degradation; direct fix if root cause clear and single-command; rollback to known-good if root cause unclear after multiple changes. Rollback is standard recovery, not failure.
Production Considerations (13)
13.1 RBAC
Separate read-only SA for daily troubleshooting (kubectl auth can-i), controlled high-privilege for destructive ops with second-person review, no long-lived cluster-admin credentials.
13.2 Change Windows & Canary
Test in staging first; production canary via maxUnavailable:0, maxSurge:1 (zero-downtime with readinessProbe); new Ingress with subset of traffic.
13.3 Required Baselines
Collect per-service Endpoints count (1m), Ingress 5xx ratio (1m), Pod restart count (1m), CoreDNS P99 latency (1m), conntrack utilization (1m), node NotReady events (real-time).
13.4 Anti-Patterns
Restart-on-issue (destroys evidence, config errors persist); manual delete Endpoints (controller recreates in seconds); delete NetworkPolicy to "unblock" (forget to restore); bypass Ingress to point Service at Pod IP (breaks on Pod recreate).
13.5 Cluster-Level Factors
Node memory pressure evicts Pods (check Conditions, Allocated); high Pod density grows iptables rules (check per-node Pod count); CNI differences (Calico/Flannel/Cilium) affect NetworkPolicy support; Service count impacts kube-proxy sync (monitor sync duration).
Complete Troubleshooting Walkthrough (14)
14.1 Symptom
14:05 alert: order service 5xx from 0.1% to 38%; users report "order page fails, refresh works".
14.2 Initial Hypothesis
"Refresh works" + 38% failure → partial backend unavailable. 3 replicas → ~33% per replica.
14.3 Commands & Findings
Pods: 2 Ready, 1 NotReady (0/1). Endpoints: only 2 IPs. EndpointSlice: third Pod ready=false. Pod not Ready but Endpoints excludes it → 38% failure not from this Pod. Describe Pod: Readiness probe failed HTTP 503. Logs: downstream inventory-api call timeout, health check marks dependency unhealthy.
14.4 Downstream Investigation
inventory-api: 1 replica, Running. Test from order Pod: curl inventory-api:8080/healthz → 200 but 4.8s latency (near timeout). inventory-api top: CPU 998m/1000m limit, memory 512Mi/512Mi limit (OOM risk).
14.5 Root Cause
inventory-api CPU limit 1 core saturated by traffic growth; latency 50ms→4.8s. order service health check depends on downstream with 3s timeout → probe fails → Pod removed from Endpoints. 38% = 33% capacity loss + some requests on remaining replicas timing out on downstream.
14.6 Fix: Stop Bleeding Then Cure
Immediate: patch inventory-api CPU limit to 4, rollout status, top pod. Long-term: scale inventory-api to 3 replicas + HPA; decouple health check from downstream dependencies (separate /readyz for readiness, /livez for liveness); add downstream timeouts/circuit breakers.
14.7 Verification
inventory-api latency 0.062s; order-api all 3 Ready; 5xx drops to 0.1% in 5 min.
14.8 Rollback Plan
If CPU still saturated after scale → rollout undo inventory-api, pivot to code optimization.
14.9 Retrospective
Timeline: 13:50 traffic peak, 13:58 inventory CPU full, 14:03 order health checks timeout, 14:05 alert, 14:12 root cause found, 14:20 scale effective. Gap: no downstream latency monitoring; CPU alert threshold too high (95% 10min) vs 5min to failure. Improvements: add downstream latency alerts, decouple health checks, HPA for critical services, lower CPU alert threshold or use "CPU > 80% of limit" predictive rule, health check timeout > 3x downstream P99. Checklist for other services: health check depends on downstream? downstream has timeout/circuit breaker? resource limits match load? single-replica risk assessed? downstream latency monitored?
Summary
Kubernetes access troubleshooting = decompose long chain into independently verifiable short chains, converge layer-by-layer from outside in. Blind Pod restart only fixes "Pod itself broken" (minority); majority are config mismatches and resource exhaustion. Key heuristics: status code → layer (404→Ingress, 503→Endpoints, 502→backend app, timeout→NetworkPolicy/kernel); empty Endpoints = top 503 cause (check first); Pod IP OK + ClusterIP fail = kube-proxy; Pod IP + ClusterIP both fail = app/Pod network; "never worked" vs "suddenly broken" = config vs change/resource; partial failure ~1/N → check N replicas for one bad; health check depending on downstream = dangerous amplification. Troubleshooting skill = knowing which layer's which metric + experience of "what normal looks like". Document everything: service, namespace, time, commands, outputs, root cause, fix — this accumulation is ops value.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Ops Community
A leading IT operations community where professionals share and grow together.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
