Operations 71 min read

Why Ping Works but HTTPS Fails: A Layered Troubleshooting Guide

This comprehensive guide explains why a server may respond to ping but fail HTTPS access, detailing a systematic layered approach covering DNS, proxy, TCP, TLS, HTTP, and application layers with concrete commands and case studies.

MaGe Linux Operations
MaGe Linux Operations
MaGe Linux Operations
Why Ping Works but HTTPS Fails: A Layered Troubleshooting Guide

Problem Background

Ping tests ICMP echo to an IP, while HTTPS involves DNS, TCP 443, proxy, TLS handshake, certificates, HTTP routing, and application response. Success in ping only confirms a small segment of the path.

Save Original Symptoms

Record timestamp, client network, target URL, proxy usage, IPv4/IPv6, failure scope, and exact error. "Cannot open" may mean DNS failure, connection timeout, connection refused, certificate error, TLS protocol mismatch, HTTP 403/502, or application exception — each requires a different fix.

# Verbose curl showing resolution, connection, and TLS
curl -v --connect-timeout 5 --max-time 15 https://api.example.com/health -o /dev/null

# Capture status code and phase timings
curl -sS -o /dev/null -w 'ip=%{remote_ip} code=%{http_code} dns=%{time_namelookup} connect=%{time_connect} tls=%{time_appconnect} first=%{time_starttransfer} total=%{time_total}
' --connect-timeout 5 --max-time 15 https://api.example.com/health
time_appconnect

is cumulative from connection start to TLS handshake completion; do not simply add individual times. Avoid -k initially to avoid masking certificate validation issues.

Step 1: Confirm DNS Points to Expected Address

Ping may resolve to IPv4 while browser prefers IPv6, or browser may use a proxy while ping goes direct. Compare resolved addresses, actual connection IP, and expected service entry.

getent ahosts api.example.com
dig +short A api.example.com
dig +short AAAA api.example.com
curl -4 -Iv --connect-timeout 5 https://api.example.com/
curl -6 -Iv --connect-timeout 5 https://api.example.com/

If -4 works but -6 times out, check AAAA record, IPv6 routing, and load balancer listeners — not the application certificate. /etc/hosts, corporate DNS, and browser secure DNS may return different answers; verify on the failing client.

Step 2: Confirm Proxy Involvement

Environment variables HTTPS_PROXY, NO_PROXY, system proxy, and browser proxy config can alter request path. Proxy auth failure often returns 407; proxy unable to reach upstream may show 502/504. Compare each systematically.

for key in HTTP_PROXY HTTPS_PROXY ALL_PROXY NO_PROXY; do
  if printenv "$key" >/dev/null; then printf '%s=set
' "$key"; fi
done
curl --noproxy '*' -Iv --connect-timeout 5 https://api.example.com/

If only proxy path fails, check proxy logs and upstream connectivity; if only browser fails while curl succeeds, compare browser proxy, HSTS, corporate certificates, and extensions — do not immediately assume server health.

Step 3: TCP 443 Connectivity

ICMP and TCP may be handled by different rules. Connection refused means a device actively rejects or port has no listener; Connection timed out often indicates path packet loss, silent firewall drop, or routing anomaly. Both need server-side and middle-device evidence.

nc -vz -w 3 api.example.com 443
ip route get 192.0.2.10
ss -lntp '( sport = :443 )'

Listening on 127.0.0.1:443 allows only local access; 0.0.0.0:443 listens on all IPv4 interfaces. IPv6 coverage must be checked separately. If entry is a load balancer, backend nodes may not listen on 443 — verify ingress architecture first.

Fixed IP Test for DNS, Preserve Hostname for TLS

When suspecting wrong DNS, use curl --resolve to map hostname to a test IP temporarily. The URL still uses the domain, preserving correct Host header and TLS SNI; direct https://IP/ often fails due to virtual host and certificate mismatch.

curl -Iv --resolve api.example.com:443:192.0.2.10 https://api.example.com/health
curl -Iv --resolve api.example.com:443:192.0.2.11 https://api.example.com/health

If node A succeeds and node B fails, prioritize backend health, certificate, and config drift. If fixed IP works but normal access fails, return to DNS/proxy/ingress scheduling.

Step 4: TLS Handshake and Certificates

When TCP succeeds but TLS fails, distinguish certificate validation, SNI, protocol version, cipher suites, mutual TLS, and handshake timeout. Expired cert, name mismatch, missing intermediate cert, or untrusted issuer all affect validation; server requiring client cert also fails at handshake.

openssl s_client -connect api.example.com:443 -servername api.example.com -showcerts </dev/null
openssl s_client -connect api.example.com:443 -servername api.example.com </dev/null 2>/dev/null | openssl x509 -noout -dates -ext subjectAltName
date -u
timedatectl status
s_client

printing a certificate does not mean validation passed; watch verify return code. Use the same trust roots as production clients. curl -k is only for controlled testing to isolate validation chain issues — never a production fix.

TLS Success May Still Be HTTP Issues

Receiving an HTTP status code means previous stages passed. 403 usually indicates auth/authorization rules; 404 points to path, virtual host, or app routing; 502/503/504 typically involve ingress-to-upstream, service unavailable, or upstream timeout — but exact meaning depends on the proxy used. Keep response headers and request IDs; check corresponding ingress logs.

curl -sS -D /tmp/https-headers.txt -o /tmp/https-body.txt https://api.example.com/health
rg -i '^(HTTP/|server:|via:|x-request-id:|location:)' /tmp/https-headers.txt

If /health works but business path fails, check auth, upstream routing, request body size, and app dependencies — do not keep changing DNS. If HTTP redirects to wrong domain or loops, verify reverse proxy scheme and Host forwarding.

MTU and Hidden Path Failures

Small ICMP packets pass while TLS handshake or large responses stall — possible path MTU, fragmentation, and ICMP unreachable messages being dropped. Also consider middlebox interference, packet loss, or client proxy issues; only pursue MTU when packet captures and size differences support it.

ping -M do -s 1400 -c 3 192.0.2.10
tcpdump -ni any host 192.0.2.10 and tcp port 443 -c 100

Capture must distinguish whether SYN sent, SYN-ACK returned, ClientHello reached server, server responded. Production troubleshooting recommends simultaneous client and server captures, correlated by 5-tuple and time; logs and captures may contain sensitive data — handle per policy.

Three Illustrative Failure Cases

Case A: IPv6 Black Hole. Ping uses IPv4, curl -4 succeeds, curl -6 times out, DNS has AAAA; fix IPv6 ingress routing or temporarily remove erroneous record, then verify from multiple networks.

Case B: Missing Certificate Chain. TCP and s_client handshake succeed, but curl reports untrusted issuer; add missing intermediate certificates per cert configuration, validate across multiple clients.

Case C: Port 443 Not Listening. ICMP succeeds, TCP returns refused; investigate reverse proxy listener and load balancer backend port mapping. These are teaching cases — a single command cannot lock down a production root cause.

A: -4 success / -6 fail → check AAAA & IPv6 path
B: handshake has cert / client validation fail → check chain, SAN, time
C: ICMP success / TCP 443 refused → check listener & ingress mapping

Change, Verification, and Rollback

Before fixing, record live DNS TTL, load balancer backends, certificate versions, and firewall rules. Change only one variable at a time (cert, DNS, firewall), validate first on test client with fixed IP, then observe real traffic success rate. Rollback prep includes old certs/configs, old DNS records, and effective times; DNS fallback still subject to cache and TTL.

getent ahosts api.example.com
curl -Iv --connect-timeout 5 https://api.example.com/health

FAQ and Phase Summary

Does ping failure guarantee HTTPS failure? No — some networks block only ICMP while TCP works.

Does curl -k success mean ready for production? No — it skips certificate validation.

Why does IP access work but domain fail? Could be DNS, SNI, Host, proxy, or certificate at any stage.

Treat "ping works" as a narrow clue. Preserve original error, start from DNS and proxy, then sequentially confirm TCP, TLS, HTTP, and application. Use controlled variable experiments to locate the faulty layer, then fix with a rollback plan.

Distinguish Connect Timeout from Read Timeout

--connect-timeout

governs connection establishment wait; --max-time limits entire transfer. If connection succeeds but read phase stalls, investigate app processing, reverse proxy upstream, and connection reuse — not DNS. Long responses or streaming endpoints may naturally take longer; set timeouts per business expectations.

curl -v --connect-timeout 3 --max-time 20 https://api.example.com/slow-endpoint -o /dev/null

Example failure record: DNS 0.02s, TCP 0.04s, TLS 0.12s, first byte 15s → network connection established, focus on ingress-to-app and backend processing. This indicates direction, not final proof; server should correlate same request ID across ingress time, forward time, and processing completion to find extra wait.

Certificate Chain, SAN, and SNI Controlled Experiments

Same IP can host multiple HTTPS sites. SNI tells server target domain during handshake; server selects certificate accordingly. Using openssl s_client -connect IP:443 without -servername may return default certificate, causing misjudgment. Certificate SAN must include accessed hostname; wildcard coverage has explicit limits.

openssl s_client -connect 192.0.2.10:443 -servername api.example.com -verify_return_error </dev/null
openssl s_client -connect 192.0.2.10:443 -servername other.example.com -verify_return_error </dev/null

If results differ, check virtual host and certificate binding. Missing intermediate certs may pass on some clients but fail on others due to different trust stores and caches; always rely on the failing client's validation result.

Troubleshooting TLS Protocol and ALPN

Client and server must agree on TLS version and cipher suites. If only old clients fail, check server minimum TLS version and client capabilities; do not enable deprecated protocols just for compatibility. HTTP/2 ALPN negotiation and app-layer behavior can also cause differences; confirm protocol version and response errors first.

curl -Iv --http1.1 https://api.example.com/health
curl -Iv --http2 https://api.example.com/health
openssl s_client -connect api.example.com:443 -servername api.example.com -alpn h2,http/1.1 </dev/null
curl --http2

requires compile-time support; missing support yields immediate tool error. HTTP/2 vs HTTP/1.1 differences point to ingress protocol translation, upstream settings, and app response — not something "ping works" can refute.

Firewall and Ingress Investigation

Client SYN sent with no response: check rule counters on client, boundary, and server simultaneously. Linux hosts may have nftables, cloud security groups, container network policies, and load balancer ACLs concurrently — checking only iptables -L is insufficient. Examine actual ingress path first, then narrow scope.

sudo nft list ruleset | rg -n '443|drop|reject'
sudo ss -lntp | rg ':443\b'
sudo tcpdump -ni any 'tcp port 443 and host 192.0.2.20' -c 50

Example uses 192.0.2.20 as client test IP. If capture shows server replied SYN-ACK but client didn't receive, check return routing and middle network; if server never saw SYN, trace upstream along ingress path. Disabling entire firewall expands attack surface and destroys original evidence — production should validate specific 5-tuple rules.

When Failure Is Browser-Specific

Curl succeeds but browser fails: check browser error codes, proxy, HSTS, cache, extensions, client certificates, and security software. Browsers auto-upgrade HTTP to HTTPS or use special DNS policies, differing from terminal commands. Compare same user, same machine, same network across incognito window and another browser; preserve devtools Network panel connection and cert info.

Comparison matrix: Browser A / Browser B / curl
Corporate network / Phone hotspot
Domain normal access / --resolve fixed correct IP
IPv4 / IPv6

Matrix changes one condition at a time, records result and timestamp. If only corporate network fails, prioritize corporate proxy and trust roots; if multiple networks fail, check domain, cert, and server.

Quick Diagnostic Record Table

| Step | Result | Evidence | Next Step |
| --- | --- | --- | --- |
| DNS A/AAAA | TBD | getent/dig | Verify ingress |
| TCP 443 | TBD | curl -v/nc | Check routing or listener |
| TLS Verify | TBD | curl/openssl | Check cert & SNI |
| HTTP Status | TBD | Status code/Request ID | Check gateway or app |

Complete records prevent handoff rework from "ping works". Post-fix acceptance must cover failing client, different networks, IPv4/IPv6, cert validation, and business paths; confirm monitoring success rate recovers and observe for a period per change window.

Layered Faults Use Same Control Commands

Run DNS, TCP, TLS, and HTTP tests from same client, same time window, leaving evidence for each layer. Command output should include errors and exit codes for handoff. A screenshot of "website won't open" cannot distinguish 403 from network timeout.

getent ahosts api.example.com
nc -vz -w 3 api.example.com 443
openssl s_client -connect api.example.com:443 -servername api.example.com -brief </dev/null
curl -sv --max-time 15 https://api.example.com/health -o /dev/null

If step 2 fails, no need to analyze cert yet; if step 3 succeeds but step 4 returns 503, focus on HTTP upstream. A later step failing does not prove all prior steps succeeded — confirm each by actual output.

Load Balancer Backend Inconsistency Symptoms

An ingress with three backends where only one has stale cert or wrong listener port yields intermittent failures; a single successful curl can mislead. Use ingress-provided backend health info and request IDs, or (with authorization) --resolve to test each backend individually. Note: if TLS terminates at LB, direct backend test uses different protocol/cert — choose correct layer per actual topology.

LB terminates TLS: Client → LB:443 → Backend:8080 HTTP
Passthrough TLS: Client → LB:443 → Backend:443 HTTPS
for i in 1 2 3; do
  curl -sS -o /dev/null -w 'code=%{http_code} ip=%{remote_ip}
' --max-time 5 https://api.example.com/health
done

Three runs are quick observation only; not guaranteed to hit each backend. Need LB routing or backend logs to confirm sample distribution.

Certificate Update Acceptance: Don't Just Check Files

New cert on disk may not be loaded by process; LB nodes may have inconsistent configs. After update, initiate real TLS handshake from failing client, verify cert serial, validity, and SAN; check all ingress instances, then monitor validation failure rate. Rollback must ensure old cert not expired.

openssl s_client -connect api.example.com:443 -servername api.example.com </dev/null 2>/dev/null | openssl x509 -noout -serial -dates -subject
curl -Iv https://api.example.com/health

Never send private key to incident channels or screenshots. Cert troubleshooting only needs public cert and metadata; private key stays within key management.

Proxy CONNECT vs Direct Connection

Accessing HTTPS via HTTP proxy: client sends CONNECT host:443 to proxy, proxy establishes upstream TCP tunnel. Proxy returning 407 means proxy auth; tunnel established but cert untrusted may indicate corporate TLS inspection vs client trust roots. curl -v shows proxy evidence but output may contain auth info — sanitize before saving logs.

curl -Iv --proxy http://proxy.example.internal:3128 https://api.example.com/health
curl -Iv --noproxy '*' https://api.example.com/health

If org policy mandates proxy, direct connect failure may be expected — don't modify firewall based on that. First verify access path requirements.

Aligning Server Logs with Client

Attach test request ID to each request; record client start time and server access log hits. If client TCP connects but no server HTTP log, check TLS termination layer; if gateway has log but app doesn't, check reverse proxy to upstream; if app logs and returns 500, check app dependencies.

curl -sv -H 'X-Debug-Request-Id: ops-test-001' https://api.example.com/health -o /dev/null

Example: Client 10:00:00 start; Gateway 10:00:00 received; App 10:00:05 received; Response 10:00:05 done. Extra 5 seconds in gateway-to-upstream segment. This is a teaching example; if log clocks unsynced, time comparison fails — prefer request ID, single-machine durations, and distributed tracing.

Post-Recovery Regression Scope

Troubleshooting often tests only the failing machine. After fix, cover: original failing client, another network segment, different DNS resolvers, IPv4/IPv6, browser and curl, business and health paths. Cert changes also need major client trust store coverage. Observe error rates and connection latency at peak to ensure fix works under load, not just low traffic.

acceptance:
  dns_expected: true
  tcp_443: success
  tls_verify: success
  http_health: 200
  business_request: success
  ipv4_ipv6: check_applicable

Record "what was fixed" and "which verification step supports root cause". Ping success is a clue, not a full HTTPS chain acceptance.

Eight Common Symptoms Quick Reference

Could not resolve host. DNS resolution phase got no target address. Check local DNS, search domains, domain spelling, proxy usage; no need to check server cert yet.

Connection refused. TCP received reject. Verify target IP correctness, 443 listener, and ingress ACL returning REJECT. Record which device emitted reject; combine with packet capture.

Connection timed out. Check routing, cloud security groups, firewall DROP, return path, and proxy; use client and ingress dual captures to locate SYN arrival.

SSL certificate problem. Cert expired, name mismatch, or trust chain issue; use SNI-enabled s_client and failing client curl to cross-check; keep cert validation in production.

Browser error, curl works. Verify browser proxy, corporate root certs, HSTS, extensions. If browser follows different redirect chain, record final URL.

Only some requests fail. Check LB backend health, cert and release version drift; use request ID to pinpoint node.

Small requests succeed, large requests stall. With evidence, check MTU, packet loss, request body limits, proxy timeouts. Don't rule out network path just because small ping works.

Returns 403/502. TLS usually established; first identify which layer generated response. 403 → auth, authorization, WAF; 502 → upstream connectivity, health, app errors. Keep response headers, request IDs, and access logs.

This quick reference only determines where to look next; must use same client, target, time, and protocol for controlled experiments to support final root cause conclusion.

Writing Shift Handoffs

Handoff record should include: impact scope, start time, failure samples, DNS results, target IP and proxy, TCP/TLS/HTTP layered evidence, changed configs, rollback method, and metrics to watch. Writing "network cause led to inaccessible" forces next shift to restart from scratch.

incident_id: example-https-01
scope: "one office network"
dns: "A expected; AAAA suspected"
tcp_ipv4: ok
tcp_ipv6: timeout
tls_ipv4: verified
http_ipv4: 200
next_action: "check IPv6 route and AAAA record"

Recorded addresses and paths must have clear owning teams; avoid throwing "client issue" to app colleagues who lack proxy/DNS visibility.

Drill: Only Some Offices Fail

Assume same https://api.example.com/health, HQ succeeds, branch times out. First determine if both locations have same DNS A/AAAA, corporate proxy, egress IP, and routing. Even if HQ ping and curl both succeed, cannot infer branch network is healthy.

getent ahosts api.example.com
curl -sv --connect-timeout 5 https://api.example.com/health -o /dev/null

Run same commands at both locations; record resolved IP, actual connected IP, proxy info, failure stage, and time. If branch uses IPv6 and HQ IPv4, force curl -4/-6 separately; if branch must use proxy, compare proxy CONNECT responses. Don't change global corporate DNS immediately; first use --resolve on single machine to verify fixed IP access.

curl -4 -Iv https://api.example.com/health
curl -6 -Iv https://api.example.com/health
curl -Iv --resolve api.example.com:443:192.0.2.10 https://api.example.com/health

If fixed IP works but normal DNS access fails, prioritize DNS or LB selection. If fixed IP also fails while branch reaches other HTTPS sites normally, check target ingress ACL, return routing, and branch egress network. Conclusion must be built on same-time controlled evidence.

Drill: Cert Rotated, Only Old Clients Fail

After new cert deployment, modern browsers work, old app reports validation failure. Check old app's trust store, whether server sends full intermediate chain, and if SAN/signature algorithm supported by old client. Don't permanently lower TLS security on entire ingress for one legacy client; evaluate client upgrade, chain completion, or separate controlled compatibility ingress.

openssl s_client -connect api.example.com:443 -servername api.example.com -showcerts </dev/null
curl -Iv --cacert /path/to/approved-ca-bundle.pem https://api.example.com/health
--cacert

must point to approved CA bundle; don't add unknown certs as roots in production trust store. Record client version, error code, cert chain, and validation results before/after change to avoid leaving dangerous "disabled validation worked" lore.

Drill: Business Path 502, Health Path 200

Health check may only confirm ingress process alive; business path calls database or other backends. For both URLs, keep response headers and request IDs; check proxy routing to which upstream, upstream connectivity, and app logs. If only large requests 502, further check request body limits, timeouts, and backend restarts.

curl -sS -D - -o /dev/null https://api.example.com/health
curl -sS -D - -o /dev/null https://api.example.com/v1/orders

Same domain root path accessible does not prove all business paths healthy. Release acceptance should include representative auth and business calls, but test accounts and data must be prepared in controlled environment.

What ICMP Success Actually Proves

ping api.example.com

typically resolves domain then sends ICMP Echo Request to one resolved IP. Echo Reply means that target at that moment replied to such packets and return path allowed echo back. It does NOT prove TCP 443 open, that same IP family is used, or that proxy, TLS, or business app are healthy. Some LB ingresses don't reply ICMP while HTTPS works, so ping failure alone cannot declare site unavailable.

More useful troubleshooting records: "DNS resolved to which IP, client actually connected to which IP, which stage failed." If domain behind CDN or multi-region scheduling, DNS results vary by location, time, resolver. Don't treat IP pinged from Office A as guaranteed backend for Office B.

Routing and Return Path Asymmetry

Client sends SYN, no response — not necessarily target host refusal; could be egress ACL, LB front firewall, routing black hole, or asymmetric return path. Client capture confirms request sent; ingress capture confirms arrival; server capture confirms reply. Three pieces of evidence must correspond to narrow fault domain. Capture within approved window, limited to target IP and packet count, avoiding unnecessary business payload collection.

Cloud environments also need security group, subnet ACL, and LB health check review. Host ss showing 443 listening doesn't mean external ingress 443 correctly forwards; conversely, backend only listening 8080 may match "LB terminates TLS then forwards HTTP" architecture. Draw actual ingress path first, then interpret listening ports.

Direct: Client → Server TCP:443 → TLS → App
LB Terminate: Client → LB TCP:443/TLS → Backend TCP:8080/HTTP
Proxy: Client → HTTP CONNECT Proxy → Target TCP:443/TLS

Stateful Connections and Idle Timeouts

Some failures appear only on long-lived connection reuse or after idle period: first request succeeds, second reuses connection already closed by middle device and fails. Browser and curl connection pool behaviors differ; single new-connection test may not reproduce. Record whether failure occurs on first connection, reused connection, mid-streaming response, or after fixed idle duration; then check proxy and LB idle timeouts.

Again, separate TCP/TLS connection success from HTTP response completion. If client waits forever while server completed response, middle proxy may buffer or lose stream; if server never received second request, check connection pool and middle network. Post-fix, test with representative connection reuse and business duration, not just a single curl -I.

HTTP Status Is Not Synonym for Network Failure

Browser seeing 401, 403, 404, or 500 means at least one HTTP response received; most basic network steps already passed. 401/403 closer to identity/authorization; 404 may come from wrong virtual host or app routing; 5xx requires finding which layer generated it. Anomalous responses may still pass through CDN cache; response headers Via, request ID, Server, and ingress logs help locate layer, but no single header is definitive.

When HTTP redirects to HTTPS or another domain (301/302), must follow final URL. Wrong redirect target can make "homepage loads, login fails" look like cert issue. Validate per business path, use test accounts for auth flows, avoid logging real cookies and auth headers in public logs.

From "Connection Timeout" to Specific Device

Client gets Connection timed out: first check DNS final resolved address and

curl -v
Trying ...

line, then record if TCP SYN sent. Client capture shows SYN sent but ingress sees nothing → fault in client-to-ingress middle path; ingress received SYN and replied SYN-ACK but client didn't get it → check return path; ingress received SYN but no reply → check ingress listener, host rules, and LB. Each judgment needs timestamp and 5-tuple; avoid comparing data from different moments.

sudo tcpdump -ni any 'host 192.0.2.10 and tcp port 443' -c 60

If target is public LB, backend hosts may not see client original SYN — LB establishes new backend connection. Capture point must be at actual TCP termination location; don't treat "backend doesn't have this IP's packets" as ingress drop.

From "Certificate Error" to Four Checks

Check hostname: URL domain in cert SAN?

Check time: cert not yet valid or expired, client clock correct?

Check chain: server sends sufficient intermediates, client trusts root?

Check handshake cert selection: same IP multi-domain, SNI correct?

Follow this order preserving evidence; better than curl -k "it opens" for actionable conclusion.

openssl s_client -connect api.example.com:443 -servername api.example.com -showcerts -verify_return_error </dev/null

TLS success may still have client cert requirement; mutual TLS scenarios need gateway config and client cert issuance confirmation. Ordinary browser without suitable client cert may show handshake or auth error; server cert is fine — don't mistakenly issue new public cert.

Capturing HTTP Redirects and Upstream Errors

curl -I

sends HEAD, not guaranteed same response as business GET/POST; some apps don't implement HEAD. For redirect confirmation, use curl -IL to see each hop; for business, use actual method and test data. Each domain in redirect chain has its own DNS, TCP, TLS stages — first page HTTPS success doesn't guarantee login redirect target success.

curl -sS -L -D /tmp/redirect.headers -o /dev/null -w 'final=%{url_effective} code=%{http_code} redirects=%{num_redirects}
' https://api.example.com/

If redirect goes to another domain, troubleshooting record includes final URL, proxy, and cert; if loop, check app and reverse proxy X-Forwarded-Proto, Host, and rewrite rules. Request/response headers may contain auth info — sanitize before sharing logs.

Client Time and Certificate Revocation

TLS validation relies on client clock. Large system time skew makes cert appear not yet valid or expired. After confirming clock, still check cert name and chain — don't attribute all validation errors to NTP. Some environments enable cert revocation checks or corporate TLS inspection; impact depends on client implementation and policy; if only corporate device path fails, consult that ingress's cert policy.

date -u
timedatectl show -p NTPSynchronized -p TimeUSec

Time service anomalies may affect other cluster systems; fix via change process to restore sync, don't manually adjust production server time arbitrarily.

IPv4, IPv6, and Happy Eyeballs

Modern clients may try IPv6 and IPv4 simultaneously, picking first to connect. Different tools, OSes, and versions have different strategies — "ping resolved to address" ≠ browser finally connects there. curl -4 vs curl -6 comparison determines each protocol stack's usability; if only IPv6 fails, check AAAA target, listener, routing, security groups, and IPv6 return path.

curl -4 -sS -o /dev/null -w '%{remote_ip} %{http_code}
' https://api.example.com/
curl -6 -sS -o /dev/null -w '%{remote_ip} %{http_code}
' https://api.example.com/

Temporarily removing bad AAAA may mitigate, but consider DNS TTL and cache; root fix is making IPv6 ingress and health checks meet promises.

Fix Evidence Chain

Production change ticket should state "Symptom — Layered Evidence — Root Cause — Change — Re-test". Example: "IPv6 timeout, IPv4 success; AAAA pointed to old LB; updated AAAA, multi-site IPv6 TLS validation and business paths restored." If only "changed firewall, seems okay", missing rule hit and original path regression, next similar fault remains hard.

Record example: original curl error; resolution results; TCP/TLS/HTTP layers;
ingress and backend logs; change object and old values; verification client and time.

Local vs Remote Troubleshooting Evidence Gap

Running curl https://localhost/ on server succeeds only proves local loopback path; external client to LB, host network, and TLS SNI path remain unverified. If server local 127.0.0.1 works but external connection refused, check if listener bound only to loopback and ingress forwarding. Conversely, external ingress TLS terminating at LB means local backend 8080 HTTP success is normal — don't force app process to listen 443 directly.

Most valuable evidence comes from the failing client . Same time, same command on healthy client as control, compare DNS, proxy, IPv4/IPv6, and cert trust roots. Env vars like HTTPS_PROXY and NO_PROXY may differ; don't close incident just because a jump box test passes.

Browser State vs Command Line Differences

Browsers maintain connection pools, DNS cache, HSTS, proxy, cookies, and cert state. Curl is independent process, may not use same system proxy or trust roots. If only browser fails, save exact error code, final URL, Network panel phases, and cert display info. Incognito window rules out some extensions/cache but not a substitute for root cause; corporate security software may uniformly inspect all browser requests.

Must compare same HTTP method and same login state. curl -I / returning 200 while browser login POST fails could be app auth, CSRF, request body, or redirect chain — not TLS. For sensitive business, avoid copying real cookies to command history; use test accounts with short-lived credentials.

Why curl -k Is Only for Diagnosis

-k

skips server cert validation, allowing expired, name-mismatched, or untrusted certs to establish connection. If default curl fails but -k succeeds, it helps narrow to validation chain, but cannot prove traffic security, nor be permanent client config. Next step must find specific issue in cert SAN, validity, trust chain, corporate proxy, or client time.

Cert errors must be judged from failing user's trust store. Server sent full chain but very old client doesn't trust root → fix may be client upgrade; server omitted intermediate cert → fix at ingress cert config. Both shouldn't be handled with same "disable validation".

Handling Faults Behind Reverse Proxy

TLS handshake and HTTP ingress normal but returns 502/504: check reverse proxy upstream address, connection errors, and timeouts. If proxy-to-upstream also uses TLS, verify proxy-as-client SNI and trust roots. Check upstream health probes — may only cover process liveness; actual business dependency on DB may still yield 5xx. Outer HTTPS request cert correct does not prove backend chain has no TLS errors.

If fault coincides with release, compare upstream address, port, protocol, timeout, and LB weights before/after change. Reproduce on one test backend first, then canary fix; don't modify multiple params at full ingress simultaneously, losing root cause confirmation.

When to Involve Network Team

When client confirms SYN sent, ingress didn't receive, or ingress replied SYN-ACK but client didn't get it — evidence points to middle path. Provide source/dest IP, port, time, capture summary, routing, and proxy info to network team. Better than "website inaccessible, please check network." If capture shows full TLS and HTTP interaction with server returning 500, app or ingress team investigates first.

Collaboration records must not contain full auth headers, PII, or private keys. On-call process should define which team owns which device level; when cloud security groups and corporate proxy both exist, clarify handoff contacts. Even if not fixed in one shift, evidence chain remains continuable.

One Layered Result Table Beats Ten "Network OK" Claims

Shift handoff must clarify each layer's test target and result. DNS tests name-to-address; TCP tests target IP 443 connectability; TLS tests SNI handshake and cert validation; HTTP tests path, Host, auth, and upstream app. Same "fail" label at different layers implicates different systems.

| Phase | Target | Result | Evidence | Next Step |
| --- | --- | --- | --- | --- |
| DNS | api.example.com | TBD | getent/dig | Check resolution |
| TCP | Target IP:443 | TBD | curl -v/nc | Check path listener |
| TLS | SNI & Cert | TBD | curl/openssl | Check handshake chain |
| HTTP | /health | TBD | Status code/Request ID | Check proxy app |

Target IP record must come from actual request, not hour-old ping address. If domain resolves to multiple addresses, record each address behavior and LB strategy; if via proxy, actual TCP connects to proxy IP — separately record proxy-to-target path.

Why Not Use curl -k to Fix Certificates

-k

skips validation, helping distinguish "handshake or app completely broken" from "cert validation failed", but simultaneously makes client unable to confirm connected service truly belongs to target domain. Writing this param permanently into app config equals actively discarding HTTPS identity verification. Production fix should target specific cause: cert expired → replace; SAN missing domain → request matching cert; incomplete chain → configure intermediates; client clock wrong → restore time sync.

If corporate private CA not yet trusted by clients, distribute CA root cert via controlled channels, not by disabling validation in every app. Save validation error text and cert chain during test; don't just write "cert issue" — same error code may hide different client trust store differences.

Same Target, Different Protocol Stacks

ping

uses ICMP, curl https uses TCP+TLS+HTTP; curl --http3 (if supported) may involve QUIC/UDP. Firewalls or LBs handle them separately. Only browser fails while curl HTTP/1.1 works → confirm browser actual protocol and fallback behavior, not demand network team use ping to prove "no problem." HTTP/3 troubleshooting must first confirm both client and server support; don't treat unsupported protocol tool error as network fault.

HTTP/2 multiplexes multiple requests; connection-layer fault impact may be more visible than single HTTP/1.1 request. Using curl --http1.1 vs curl --http2 for comparison, keep URL, proxy, auth identical; record actual negotiated protocol. Protocol differences point to ingress and app adaptation, not mandatory downgrade.

Firewall Rule Fix Scope Control

Seeing 443 rejected, don't shut down entire inbound firewall. First check source CIDR, dest IP, port, protocol; confirm which rule counter increments; then allow per minimal access need. If cloud security group and host nftables both exist, both layers must allow, but which layer to change depends on evidence. Pre-change export rules, record affected business; post-change verify allowed target requests AND that other ports remain closed.

sudo nft -a list ruleset | rg -n '443|drop|reject'

In container or LB-terminated TLS architectures, host 443 rules may not be final control point. Packet capture and ingress topology must precede allow operations to avoid opening wrong place.

Preventing Recurrence

External domain monitoring should validate TCP/TLS/HTTP from multiple networks and protocol families, not just ping domain. Cert expiry alerts must check ingress actual served cert, not just disk file. Multi-backend LB needs per-node health and config drift checks. Post-release, use real business paths not just /health for controlled synthetic probes.

synthetic_checks:
- name: https-ipv4
  host: api.example.com
  family: ipv4
  assert: [tls_valid, http_200]
- name: https-ipv6
  host: api.example.com
  family: ipv6
  assert: [tls_valid, http_200]

Above config describes monitoring design, not ready tool syntax; IPv6 check needed only if domain publishes AAAA. Probe failures should retain phase timings and target IP so on-call can start from failing layer.

Full Layered Troubleshooting Drill Record

Assume 3 PM customer reports payment interface down. On-call tests from failing user's same network: getent returns expected IPv4 and IPv6. curl -4 TCP/TLS/HTTP all succeed. curl -6 times out at TCP connect; browser errors also concentrated on IPv6-capable terminals.

Form hypothesis "IPv6 ingress path abnormal" rather than immediately declaring "app completely fine." Next: check AAAA, IPv6 LB listeners and security groups, ingress captures, backend health.

curl -4 -sv --connect-timeout 5 https://api.example.com/health -o /dev/null
curl -6 -sv --connect-timeout 5 https://api.example.com/health -o /dev/null

If fixed correct IPv6 via --resolve works but normal access fails, prioritize DNS AAAA returning old ingress; if fixed IP also fails, check routing, ACL, return path. Mitigation: within authorized change process, remove bad AAAA or fix ingress, then re-test in failing network, other networks, and actual business paths. DNS cache and TTL cause async recovery; observation period shouldn't close on single machine success.

15:00 customer report; 15:05 obtain curl error & target IP;
15:10 IPv4 success, IPv6 TCP timeout; 15:20 locate wrong ingress;
15:30 fix; 15:40 multi-site IPv6 TLS/HTTP verification done.
This timeline is purely hypothetical for teaching.

If curl -6 runs on test machine without IPv6, failure doesn't prove target ingress fault. Cross-check by accessing another confirmed IPv6 site, or run from failing client with confirmed IPv6 egress. Every tool conclusion has applicability conditions.

Troubleshooting Handbook Must Not Omit Limitations

ping

may be disabled. nc doesn't check cert or HTTP. openssl s_client output cert ≠ validation passed. curl -I uses HEAD, may differ from real GET/POST. --resolve must keep URL domain.

Packet capture only valid at chosen collection points and paths.

On-call should write these limits in records; don't extend single tool result to whole chain conclusion.

| Command | Can Prove | Cannot Prove |
| --- | --- | --- |
| ping | ICMP echo reply at address | TCP/TLS/HTTP normal |
| nc -vz | TCP connectable | Cert & app normal |
| s_client | TLS handshake details | HTTP business path healthy |
| curl -v | One real URL request | All nodes and users normal |

Post-fix results also need same client, same URL, same business method re-test. If fault was in browser login flow, final must use controlled test account to complete login, not just test GET /health.

Monitoring and Ticket Closure Loop

Monitor DNS error rate, TCP connection failures, TLS validation failures, and HTTP 5xx separately — each alert points to different troubleshooting entry. Synthetic probes from multiple locations with target IP and phase timings. Cert expiry alerts against ingress actual served cert, avoiding disk file updated but proxy process still using old cert. Incident ticket finally records fix object, rollback path, re-test locations, and observation window.

When frontline receives "ping works but can't open", can immediately assign to specific layer instead of defaulting network, app, and security teams asking each other in a group. Troubleshooting speed comes from credible controlled experiments and complete evidence, not more undifferentiated retries.

Final Verification: Return to Original Failure Path

Whether fix is cert, DNS, firewall, or proxy, ultimately must re-request from original failing location, terminal, and business URL with full cert validation. Fixed IP --resolve test success only proves chosen address can serve; normal domain resolution and LB may still route users to bad nodes. Observe success rates across different networks and backends for a period to confirm fix covers original impact scope.

If fault originally in login or payment path, also use controlled test account to walk full redirect and business requests, checking Cookie, Host, proxy, and upstream timeouts. Health check 200 only proves health check reachable; if cross-domain redirect uses another cert or LB, second chain must be included in acceptance.

If service behind CDN, edge nodes differ by geography and cache state. Compare client's actual edge IP, response header request ID, and origin errors; one edge 502 doesn't mean all origin nodes faulty. Fixed IP troubleshooting must first confirm it belongs to current valid ingress, avoiding stale nodes or doc example addresses as real targets.

External monitoring points should retain final target IP, TLS validation result, and HTTP status. Next alert, on-call can start directly from failing layer without re-guessing with ping.

Record observation window and rollback contacts post-fix; ensure future cert renewals, DNS changes, and proxy releases have corresponding layered acceptance.

Evidence must be reproducible for conclusions to be reliable.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

IPv6curlopensslload balancerDNS resolutionHTTPS troubleshootingproxy configurationnetwork diagnosticscertificate validationTLS debugging
MaGe Linux Operations
Written by

MaGe Linux Operations

Founded in 2009, MaGe Education is a top Chinese high‑end IT training brand. Its graduates earn 12K+ RMB salaries, and the school has trained tens of thousands of students. It offers high‑pay courses in Linux cloud operations, Python full‑stack, automation, data analysis, AI, and Go high‑concurrency architecture. Thanks to quality courses and a solid reputation, it has talent partnerships with numerous internet firms.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.