NetEase Interview: Debugging 900% CPU Spikes in MySQL and Java Processes

This article details systematic troubleshooting for 900% CPU spikes in MySQL and Java processes, covering top/jstack diagnostics, show processlist analysis, index optimization, cache integration, and fixing busy-wait loops with BlockingQueue.take(), illustrated through real production case studies.

Architect's Guide
Architect's Guide
Architect's Guide
NetEase Interview: Debugging 900% CPU Spikes in MySQL and Java Processes

Scenario 1: MySQL Process CPU Spikes to 900%

High concurrency combined with poorly performing SQL (e.g., missing indexes) causes rapid CPU spikes. Enabling slow query logging under such conditions worsens performance.

Troubleshooting Steps

Use top to confirm mysqld is the culprit.

Run show processlist; to inspect active sessions and identify resource-intensive queries.

For each high-cost SQL, examine the execution plan ( EXPLAIN) to check for missing indexes or excessive data volume.

Resolution Process

Kill offending threads ( kill <thread_id>) while monitoring CPU drop.

Apply fixes: add missing indexes, rewrite SQL, or tune memory parameters.

If connection count surges, coordinate with application team to limit connections.

Optimize iteratively: apply one change, observe, then proceed.

Real-World MySQL Case Study

Production MySQL CPU hit 900%+. Data volume was only a few million rows. Investigation revealed:

Many queries bypassed cache and hit MySQL directly. show processlist; showed numerous identical queries stuck in query state:

select id from user where user_code = 'xxxxx';
show index from user;

confirmed user_code lacked an index. Adding the index allowed normal execution.

Soon after, request timeouts spiked again. Root cause: slow query log was enabled; massive log writes degraded performance. Disabling it dropped CPU to ~300%.

Migrating real-time queries to Redis cache brought CPU to 70–80%.

Key Takeaways:

Avoid enabling slow query log during severe CPU saturation — disk I/O from logging worsens the problem. show processlist directly pinpoints problematic SQL; analyze for index, lock, wide-column, or large-table issues.

Always use a cache layer (Redis) to reduce MySQL query frequency.

Memory tuning is another viable lever.

Scenario 2: Java Process CPU Spikes to 900%

Java processes normally run at 100–200% CPU. Spikes to 900% typically indicate infinite loops or excessive GC under high concurrency.

Troubleshooting Steps

top

→ identify high-CPU Java process PID. top -Hp <PID> → list threads; sort by CPU (Shift+P) to find hot threads.

Convert thread PID to hex: printf "%x\n" <TID> (e.g., 30309 → 0x7665).

Capture thread dump: jstack -l <PID> > ./jstack_result.txt.

Search dump for hex nid: cat jstack_result.txt | grep -A 100 7665 or jstack <pid> | grep -A 200 <nid> to locate the exact method.

Analyze the code at that stack frame.

Common Causes & Fixes

Empty spin loop: Use Thread.sleep() or locking to block the thread.

Massive object allocation in loop (e.g., loading 1M+ rows from MySQL): Reduce allocations or introduce object pooling.

Selector busy-wait (Netty): Rebuild selector after a threshold of empty selects, per Netty's approach.

Real-World Java Case Study: 700% CPU Spike

A newly deployed service drove physical servers to frequent crashes. Diagnosis: top showed PID 29706 at 700%+ CPU. top -Hp 29706 revealed multiple threads at 90%+ CPU; picked TID 30309. printf "%x\n" 30309 → 0x7665. jstack -l 29706 > ./jstack_result.txt then grep -A 100 7665 pointed to ImageConverter.run().

Problematic code used a LinkedBlockingQueue with poll() in a tight loop:

while (isRunning) {
    if (dataQueue.isEmpty()) {
        continue; // busy-wait when queue empty
    }
    byte[] buffer = dataQueue.poll();
    // process buffer
}

When the queue stayed empty, continue caused a CPU-hogging spin loop. LinkedBlockingQueue offers two retrieval methods:

// Blocks until element available
E take() throws InterruptedException;
// Returns null immediately if empty
E poll();

Fix: Replace poll() + empty-check with blocking take():

while (isRunning) {
    byte[] buffer = new byte[0];
    try {
        buffer = dataQueue.take(); // blocks until data arrives
    } catch (InterruptedException e) {
        e.printStackTrace();
    }
    // process buffer
}

After restart, process CPU dropped below 10% and remained stable.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

production debuggingdatabase indexingjstackCPU troubleshootingblocking queuecaching optimizationJava thread analysisMySQL performance tuning
Architect's Guide
Written by

Architect's Guide

Dedicated to sharing programmer-architect skills—Java backend, system, microservice, and distributed architectures—to help you become a senior architect.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.