Troubleshooting CPU Spikes: From System Metrics to Code-Level Root Cause Analysis
A comprehensive guide to diagnosing CPU spikes on Linux servers, covering system-level analysis with top/vmstat, process/thread identification via ps/pidstat, stack tracing with jstack/perf/py-spy, container/Kubernetes resource limits, and real-world case studies including regex bottlenecks, cron job spikes, and CPU throttling.
Problem Background
At 2 AM, an alert triggers: an application server's CPU usage exceeds 90% for 15 minutes. The on-call engineer must answer three questions within five minutes: which process consumes CPU, which thread and code path inside that process, and whether the spike stems from normal load growth or engineering defects like deadlocks, lock contention, or GC anomalies. Restarting only masks symptoms; root cause analysis prevents recurrence.
Applicable Scenarios
Physical/virtual Linux servers with sustained high CPU triggering alerts
Cloud instances (ECS, EC2) needing distinction between resource shortage and application issues
Containerized deployments (Docker, Kubernetes) requiring in-container process identification
Java, Go, Python, Node.js backends needing thread/stack-level drill-down
Differentiating "high CPU but responsive" from "high CPU with visible slowdown"
Not covered: memory (OOM/leaks), disk I/O bottlenecks, or network issues — though cross-references are noted.
Core Concepts
3.1 CPU Usage vs Load Average
Load average reflects runnable + uninterruptible-sleep (D state) processes. High load may indicate CPU saturation or disk I/O waits. Use vmstat 1 or top to check %wa (I/O wait). If %wa is high while %us + %sy are low, the bottleneck is disk I/O, not CPU compute.
3.2 User vs Kernel CPU
%us: user-space computation (business logic, regex, serialization) %sy: kernel-space (syscalls, context switches, lock arbitration)
High %sy (>30%) suggests excessive syscalls, context switches, or network packet processing (soft interrupts %si). High %us points to application code needing thread/stack analysis.
3.3 Single-Core vs Multi-Core Saturation
An 8-core machine showing 100% total CPU could mean all cores at ~12.5% (normal concurrency) or one core at 100% (single-thread hotspot). Press 1 in top to view per-core usage — a critical but often skipped step.
3.4 Process States (R/S/D/Z/T)
R: Running/runnable S: Interruptible sleep (waiting for network) D: Uninterruptible sleep (disk I/O) — kill -9 ineffective Z: Zombie (parent hasn't wait()) T: Stopped by signal
Many D processes indicate I/O issues; switch to iostat / iotop.
3.5 Context Switches & Soft Interrupts
vmstat 1shows cs (context switches/sec). Abnormally high values (e.g., baseline thousands → hundreds of thousands) indicate thread thrashing from misconfigured pools, lock contention, or concurrent cron jobs. %si reflects soft interrupt CPU from network bursts, small packets, or uneven NIC interrupt distribution.
Overall Investigation Flow
System-wide confirmation : top / vmstat to distinguish us / sy / wa and per-core balance
Locate process : ps aux --sort=-%cpu, pidstat for top PID
Locate thread : top -H -p <pid> or ps -eLf for top TID
Locate call stack : jstack (Java), pprof (Go), perf / strace (generic)
Correlate with business logs & deployments : traffic spikes, new releases, slow SQL, retry storms
Classify root cause : capacity vs code defect vs config vs external dependency
Apply fix : rate-limit, scale, restart, patch
Verify : metrics return to baseline, no new anomalies
Postmortem : archive symptoms, root cause, fix, screenshots
Principle: macro → micro, system → process → thread → code. Each step narrows scope with evidence; no jumping to conclusions.
Hands-On Walkthrough: E-Commerce Order Service
Step 1: Overall Load
uptime
# 15:32:07 up 45 days, 3:12, 2 users, load average: 12.45, 9.30, 6.808-core machine, 1-min load (12.45) > cores → queuing. Rising 1-min vs 5/15-min indicates recent spike.
Step 2: CPU vs I/O
top
# %Cpu(s): 78.3 us, 8.1 sy, 0.0 ni, 10.2 id, 0.5 wa, 0.0 hi, 2.9 si, 0.0 st us=78.3%, wa=0.5% → user-space compute bottleneck, not I/O. Per-core view ( 1 key) shows multiple cores high → likely traffic increase, but verify at process level.
Step 3: Identify Process
ps aux --sort=-%cpu | head -n 15
# USER PID %CPU %MEM VSZ RSS TTY STAT START TIME COMMAND
# app 8823 210.5 4.2 2483920 689120 ? Sl 14:58 12:33 java -jar order-service.jar
# app 8824 15.3 3.8 2210400 601200 ? Sl 14:58 1:02 java -jar payment-service.jarPID 8823 at 210.5% (multi-threaded, >2 cores). pidstat -u 1 5 shows sustained vs burst pattern.
Step 4: Identify Thread
top -H -p 8823
# PID USER %CPU %MEM TIME+ COMMAND
# 8901 app 95.2 4.2 8:12.33 java
# 8902 app 93.8 4.2 8:05.11 java
# 8823 app 1.1 4.2 0:02.44 javaThreads 8901, 8902 near 100% each — hotspot threads. Convert TID to hex: printf '%x\n' 8901 → 2295.
Step 5: Java Stack Trace ( jstack )
jstack 8823 > /tmp/jstack_8823_$(date +%H%M%S).log
grep -A 30 "nid=0x2295" /tmp/jstack_8823_143205.logOutput shows thread RUNNABLE at PromotionMatcher.matchRules(PromotionMatcher.java:88) → calculate → applyDiscount → createOrder. Three consecutive samples stuck at same line confirms hotspot (e.g., nested loops, regex backtracking, inefficient collection traversal).
Step 6: Non-Java Generic Tools
perf : perf top -p 8823 reveals regexp_match_internal at 42% — same regex bottleneck. perf record -p 8823 -g -- sleep 30; perf report for call graphs.
strace : strace -c -p 8823 -f summarizes syscall time/count (e.g., frequent read / write / futex). Warning: strace slows target significantly; use timeout 10 in production.
Step 7: Cross-Verify with Logs & Deployments
grep "createOrder" /app/logs/order-service.log | grep "2026-09-11 14:5" | wc -l
ls -la /app/order-service/order-service.jar
git log --since="2 days ago" --oneline -- src/main/java/com/example/order/util/PromotionMatcher.javaIf PromotionMatcher.java changed pre-promotion and rule count surged, root cause chain: new O(n²) logic + traffic + rule growth = threshold breach.
Language-Specific Supplements
Go: pprof
Import net/http/pprof on internal port (e.g., 127.0.0.1:6060). Collect:
curl -s http://127.0.0.1:6060/debug/pprof/profile?seconds=30 -o /tmp/cpu_$(date +%H%M%S).prof. Analyze: go tool pprof -top /tmp/cpu_143205.prof shows regexp.(*Regexp).doExecute at 29.5% flat. Check goroutine leaks:
curl -s http://127.0.0.1:6060/debug/pprof/goroutine?debug=1 | head -n 5.
Python: py-spy
No code changes needed. pip install py-spy; py-spy dump --pid 8823 or py-spy top --pid 8823. Requires CAP_SYS_PTRACE; coordinate with security in containerized envs.
Container & Kubernetes Specifics
Map Container to Host PID
docker inspect --format '{{.State.Pid}}' order-serviceThen use host PID with top -H, jstack, perf, py-spy.
K8s Pod-Level CPU
kubectl -n production top pod
kubectl -n production top pod order-service-7d9f8c6b4-xk2p9 --containersRequires metrics-server. Check kubectl -n kube-system get pods -l k8s-app=metrics-server.
Debug Distroless Containers
kubectl -n production debug order-service-7d9f8c6b4-xk2p9 -it --image=busybox --target=order-serviceCreates ephemeral debug container with tools; no image rebuild.
CPU Throttling Detection
kubectl -n production exec -it order-service-7d9f8c6b4-xk2p9 -- cat /sys/fs/cgroup/cpu/cpu.stat
# nr_periods 921600
# nr_throttled 45832
# throttled_time 128374920123 nr_throttled/nr_periods ≈ 5%and growing throttled_time = throttling. Fix: raise limits.cpu from 500m to 2, requests.cpu to 1. Verify nr_throttled stops growing; P99 latency drops from 850ms to 180ms.
Cloud Server Nuances
CPU credits : Burstable instances (AWS T-series) throttle at baseline when credits exhausted — console shows credit balance, not local top.
Shared/entry tiers : Not designed for sustained load; upgrade to compute-optimized.
Cross-AZ latency : Remote DB/cache increases connection hold time → thread pool exhaustion → indirect CPU rise from retries.
Common Misdiagnoses
Symptom : High load average → Wrong Conclusion : CPU bottleneck → Correct Approach : Check %wa and D state processes
Symptom : Process %CPU > 100% → Wrong Conclusion : Process anomaly → Correct Approach : Normal for multi-threaded on multi-core
Symptom : Container top shows host CPU → Wrong Conclusion : Container resource pressure → Correct Approach : Use cpu.stat or kubectl top Symptom : Low CPU but slow response → Wrong Conclusion : CPU not the issue → Correct Approach : Check cloud credits, cgroup throttling, GC pauses, lock waits
Symptom : Single jstack shows thread in method → Wrong Conclusion : Dead loop → Correct Approach : Need multiple samples; sustained stall is key
Tool Permissions & Security
Restrict jstack, perf, strace, py-spy to designated SRE accounts
Clean temp files ( /tmp) containing stack traces with potential PII
Evaluate ptrace (used by strace / py-spy) impact on namespace isolation in multi-tenant clusters
Automated Snapshot Script
Bash script cpu_snapshot.sh collects: uptime, vmstat, mpstat, top processes, memory/IO overview, and if PID provided, thread dump + jstack. Outputs timestamped tarball. Safe (read-only), robust (handles missing tools), version-controlled in team scripts/ repo.
Additional Case Studies
Case 2: Periodic Cron Spike
Hourly CPU spikes (3-5 min) with timeouts. crontab -l reveals 0 * * * * /app/scripts/sync_inventory.sh. Logs show duration grew from 20s to 4+ min as SKUs grew 20k→300k. pidstat -u 1 300 confirms spike alignment. Root cause: per-row DB queries in loop. Fix: batch 500 rows per SQL, reduce frequency to 6-hourly. Post-fix: 40s runtime, CPU spike 90%→30%, timeouts gone.
Case 3: Container Throttling "False Slowness"
K8s migration → P99 latency up, but kubectl top shows 40% CPU. cpu.stat reveals throttling (5% periods throttled, 128s cumulative). Deployment had limits.cpu=500m; bursty parallel sub-tasks (inventory, promo, messaging) exceeded quota. Fix: limits.cpu=2, requests.cpu=1. Throttling stops, P99 850ms→180ms.
Case 4: DB Connection Pool Exhaustion → Indirect CPU Rise
CPU 30%→75%, latency up, timeouts. Threads show distributed moderate usage (not hotspot). Logs: Could not get JDBC Connection. MySQL Threads_connected at max; HikariCP maximumPoolSize=20 saturated. New report query: full scan + filesort (tens of thousands rows, seconds each). EXPLAIN shows type=ALL, Using filesort. Fix: add index on orders.created_at; consider read-replica routing. Post-fix: query ms→ms, pool 5-8 active, CPU back to 30%.
Case 5: Nginx Worker Misconfiguration
Worker processes 40-70% CPU, static assets slow. nginx -T shows worker_processes 2 on 8-core box + gzip_comp_level 9. Fix: worker_processes auto, gzip_comp_level 5, gzip_min_length 1024, worker_connections 4096, epoll, multi_accept on. Reload: nginx -t && nginx -s reload. Workers now 8, CPU per worker single-digit, static latency normal.
Key Terminology Reference
us— User CPU Time — App compute sy — System CPU Time — Syscalls/scheduling wa — IO Wait — Disk wait si — Software Interrupt — Network packets LWP — Light Weight Process — Linux thread TID — Thread ID — Thread identifier cs — Context Switch — Switch count/sec GC — Garbage Collection — Auto memory mgmt GIL — Global Interpreter Lock — Python thread limit cgroup — Control Group — Linux resource isolation Throttling — CPU limit enforcement — Container quota exceed Burst — Burst credits — Cloud burstable CPU
Post-Incident Actions
Fill monitoring gaps (e.g., cpu.stat throttling, connection pool metrics)
Add scenario to regular load tests (e.g., promo rule surge)
Codify anti-patterns in code review checklist (inefficient regex, missing indexes, row-by-row processing)
Recalculate capacity with real peak data and growth projections
Transform firefighting into fire prevention.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
MaGe Linux Operations
Founded in 2009, MaGe Education is a top Chinese high‑end IT training brand. Its graduates earn 12K+ RMB salaries, and the school has trained tens of thousands of students. It offers high‑pay courses in Linux cloud operations, Python full‑stack, automation, data analysis, AI, and Go high‑concurrency architecture. Thanks to quality courses and a solid reputation, it has talent partnerships with numerous internet firms.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
