Production Services Randomly Dropping Offline: A Week-Long Debugging Journey to a Linux Kernel Bug
An engineer details a week-long investigation into random microservice disappearances from Nacos in a Spring Cloud Alibaba cluster, systematically ruling out memory, CPU, disk, network, Nacos server/client issues, and JVM problems before discovering a Linux kernel bug causing JVM pauses, resolved by kernel upgrade.
Background
The production environment runs on 11 Alibaba Cloud servers using Spring Cloud Alibaba to build a microservice cluster of 60 services, all registered to a single Nacos cluster. Traffic flows through nginx → Spring Cloud Gateway → business microservices. Key versions: Spring Boot 2.2.5.RELEASE, Spring Cloud Hoxton.SR3, Spring Cloud Alibaba 2.2.1.RELEASE, Java 1.8.
Incident
During a holiday, the team received alerts that the gateway could not find services. The Nacos console showed individual services missing (deregistered). The issue recurred every few days at random times, with different services dropping each time. Manual kill+restart of affected services restored stability for 2–3 days.
Investigation and Resolution
Suspect 1: Server Memory Exhaustion
Alibaba Cloud console showed missing metrics (memory, CPU, load) for the faulty machines. Running free -m on the servers showed normal memory usage. An Alibaba Cloud ticket was opened to fix the console display; after the fix, the dropout issue persisted.
Suspect 2: CPU Saturation
Commands felt responsive; top showed normal CPU usage.
Suspect 3: Disk Full
du -sh *revealed ample disk space.
Suspect 4: Network Issues
Since service dropout implies missed heartbeats, network became the prime suspect. Commands telnet, mtr -n, netstat -nat | grep "TIME_WAIT" | wc -l gave only a rough view. Enabled tcp_tw_reuse via echo 1 > /proc/sys/net/ipv4/tcp_tw_reuse to reuse TIME_WAIT sockets. Nacos client logs showed no records of the dropout events.
Suspect 5: Nacos Server Failure
Checked the Nacos cluster servers' basic metrics (memory, CPU, disk) — no anomalies, as dozens of other services ran normally.
Nacos server logs confirmed active deregistration of the missing services. Yet some services on the same machine remained healthy, raising the question of why only a few were randomly deregistered.
Suspect 6: Microservice Resource Overuse
Deployment scripts and load balancing were identical across machines, and the same services ran fine elsewhere. Increased heap size for each microservice and added stack-trace printing; after waiting, the issue recurred but no stack traces appeared because the processes did not exit.
Nacos Client Version Check
Found a blog post noting that the Nacos client version in use (1.4.1, bundled with Spring Cloud Alibaba 2.2.1) had a known heartbeat issue. Upgraded the client; after several days the problem persisted, so focus returned to the microservices themselves.
Arthas Monitoring and Heartbeat Analysis
Used Arthas to monitor metrics — all healthy services looked normal. Hypothesized that the heartbeat thread might be killed. Examined Nacos client source code (1.x heartbeat function) and used Arthas to watch heartbeat packets. When the failure recurred, Arthas hung with no output.
Custom Dropout Listener
Built a program to detect service dropouts in real time during working hours. Finally captured an event during business hours and immediately logged into the server.
JVM State Investigation
Basic server metrics were normal. Arthas could not attach to the process. jstat showed normal GC. jmap / jstack failed to attach; used jstack -F -l 25944 > heap.txt (forced, less info). The 10,000+ line stack dump revealed no deadlock or obvious anomaly. Meanwhile, the listener reported the services had auto-recovered, making the issue harder to catch.
Reproduction and JVM Pause Hypothesis
On the second occurrence, the same pattern appeared: Arthas unable to attach → jstack -F → services recovered. Hypothesized that the JVM was experiencing a "stop-the-world" pause (fake death), freezing all threads including heartbeat and logging.
Linux Kernel Bug Discovery
Searched for "JVM pause troubleshooting" and found a credible answer pointing to a Linux kernel bug. Compared kernel versions across machines with uname -r. The faulty machine ran an older kernel (3.10.0-957.el7.x86_64) while healthy machines ran a newer version (3.10.0-1160.el7.x86_64).
Resolution
Submitted an Alibaba Cloud ticket to upgrade the kernel. After the upgrade and reboot, the issue disappeared over two days of observation. The root cause was a Linux kernel bug triggering JVM pauses, which caused heartbeat loss and Nacos deregistration.
Closing Thoughts
The issue tormented the team for over a week. The debugging process was a mix of frustration and breakthroughs. As engineers, maintaining curiosity and continuous learning is essential.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Architect's Guide
Dedicated to sharing programmer-architect skills—Java backend, system, microservice, and distributed architectures—to help you become a senior architect.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
