Operations 65 min read

Troubleshooting CPU Spikes: From System Metrics to Code-Level Root Cause Analysis

A comprehensive guide to diagnosing CPU spikes on Linux servers, covering system-level analysis with top/vmstat, process/thread identification via ps/pidstat, stack tracing with jstack/perf/py-spy, container/Kubernetes resource limits, and real-world case studies including regex bottlenecks, cron job spikes, and CPU throttling.

MaGe Linux Operations
MaGe Linux Operations
MaGe Linux Operations
Troubleshooting CPU Spikes: From System Metrics to Code-Level Root Cause Analysis

Problem Background

At 2 AM, an alert triggers: an application server's CPU usage exceeds 90% for 15 minutes. The on-call engineer must answer three questions within five minutes: which process consumes CPU, which thread and code path inside that process, and whether the spike stems from normal load growth or engineering defects like deadlocks, lock contention, or GC anomalies. Restarting only masks symptoms; root cause analysis prevents recurrence.

Applicable Scenarios

Physical/virtual Linux servers with sustained high CPU triggering alerts

Cloud instances (ECS, EC2) needing distinction between resource shortage and application issues

Containerized deployments (Docker, Kubernetes) requiring in-container process identification

Java, Go, Python, Node.js backends needing thread/stack-level drill-down

Differentiating "high CPU but responsive" from "high CPU with visible slowdown"

Not covered: memory (OOM/leaks), disk I/O bottlenecks, or network issues — though cross-references are noted.

Core Concepts

3.1 CPU Usage vs Load Average

Load average reflects runnable + uninterruptible-sleep (D state) processes. High load may indicate CPU saturation or disk I/O waits. Use vmstat 1 or top to check %wa (I/O wait). If %wa is high while %us + %sy are low, the bottleneck is disk I/O, not CPU compute.

3.2 User vs Kernel CPU

%us

: user-space computation (business logic, regex, serialization) %sy: kernel-space (syscalls, context switches, lock arbitration)

High %sy (>30%) suggests excessive syscalls, context switches, or network packet processing (soft interrupts %si). High %us points to application code needing thread/stack analysis.

3.3 Single-Core vs Multi-Core Saturation

An 8-core machine showing 100% total CPU could mean all cores at ~12.5% (normal concurrency) or one core at 100% (single-thread hotspot). Press 1 in top to view per-core usage — a critical but often skipped step.

3.4 Process States (R/S/D/Z/T)

R

: Running/runnable S: Interruptible sleep (waiting for network) D: Uninterruptible sleep (disk I/O) — kill -9 ineffective Z: Zombie (parent hasn't wait()) T: Stopped by signal

Many D processes indicate I/O issues; switch to iostat / iotop.

3.5 Context Switches & Soft Interrupts

vmstat 1

shows cs (context switches/sec). Abnormally high values (e.g., baseline thousands → hundreds of thousands) indicate thread thrashing from misconfigured pools, lock contention, or concurrent cron jobs. %si reflects soft interrupt CPU from network bursts, small packets, or uneven NIC interrupt distribution.

Overall Investigation Flow

System-wide confirmation : top / vmstat to distinguish us / sy / wa and per-core balance

Locate process : ps aux --sort=-%cpu, pidstat for top PID

Locate thread : top -H -p <pid> or ps -eLf for top TID

Locate call stack : jstack (Java), pprof (Go), perf / strace (generic)

Correlate with business logs & deployments : traffic spikes, new releases, slow SQL, retry storms

Classify root cause : capacity vs code defect vs config vs external dependency

Apply fix : rate-limit, scale, restart, patch

Verify : metrics return to baseline, no new anomalies

Postmortem : archive symptoms, root cause, fix, screenshots

Principle: macro → micro, system → process → thread → code. Each step narrows scope with evidence; no jumping to conclusions.

Hands-On Walkthrough: E-Commerce Order Service

Step 1: Overall Load

uptime
# 15:32:07 up 45 days, 3:12, 2 users, load average: 12.45, 9.30, 6.80

8-core machine, 1-min load (12.45) > cores → queuing. Rising 1-min vs 5/15-min indicates recent spike.

Step 2: CPU vs I/O

top
# %Cpu(s): 78.3 us, 8.1 sy, 0.0 ni, 10.2 id, 0.5 wa, 0.0 hi, 2.9 si, 0.0 st
us=78.3%

, wa=0.5% → user-space compute bottleneck, not I/O. Per-core view ( 1 key) shows multiple cores high → likely traffic increase, but verify at process level.

Step 3: Identify Process

ps aux --sort=-%cpu | head -n 15
# USER  PID %CPU %MEM  VSZ   RSS TTY STAT START TIME COMMAND
# app  8823 210.5 4.2 2483920 689120 ? Sl 14:58 12:33 java -jar order-service.jar
# app  8824 15.3 3.8 2210400 601200 ? Sl 14:58 1:02 java -jar payment-service.jar

PID 8823 at 210.5% (multi-threaded, >2 cores). pidstat -u 1 5 shows sustained vs burst pattern.

Step 4: Identify Thread

top -H -p 8823
# PID USER %CPU %MEM TIME+ COMMAND
# 8901 app 95.2 4.2 8:12.33 java
# 8902 app 93.8 4.2 8:05.11 java
# 8823 app 1.1 4.2 0:02.44 java

Threads 8901, 8902 near 100% each — hotspot threads. Convert TID to hex: printf '%x\n' 8901 → 2295.

Step 5: Java Stack Trace ( jstack )

jstack 8823 > /tmp/jstack_8823_$(date +%H%M%S).log
grep -A 30 "nid=0x2295" /tmp/jstack_8823_143205.log

Output shows thread RUNNABLE at PromotionMatcher.matchRules(PromotionMatcher.java:88) → calculate → applyDiscount → createOrder. Three consecutive samples stuck at same line confirms hotspot (e.g., nested loops, regex backtracking, inefficient collection traversal).

Step 6: Non-Java Generic Tools

perf : perf top -p 8823 reveals regexp_match_internal at 42% — same regex bottleneck. perf record -p 8823 -g -- sleep 30; perf report for call graphs.

strace : strace -c -p 8823 -f summarizes syscall time/count (e.g., frequent read / write / futex). Warning: strace slows target significantly; use timeout 10 in production.

Step 7: Cross-Verify with Logs & Deployments

grep "createOrder" /app/logs/order-service.log | grep "2026-09-11 14:5" | wc -l
ls -la /app/order-service/order-service.jar
git log --since="2 days ago" --oneline -- src/main/java/com/example/order/util/PromotionMatcher.java

If PromotionMatcher.java changed pre-promotion and rule count surged, root cause chain: new O(n²) logic + traffic + rule growth = threshold breach.

Language-Specific Supplements

Go: pprof

Import net/http/pprof on internal port (e.g., 127.0.0.1:6060). Collect:

curl -s http://127.0.0.1:6060/debug/pprof/profile?seconds=30 -o /tmp/cpu_$(date +%H%M%S).prof

. Analyze: go tool pprof -top /tmp/cpu_143205.prof shows regexp.(*Regexp).doExecute at 29.5% flat. Check goroutine leaks:

curl -s http://127.0.0.1:6060/debug/pprof/goroutine?debug=1 | head -n 5

.

Python: py-spy

No code changes needed. pip install py-spy; py-spy dump --pid 8823 or py-spy top --pid 8823. Requires CAP_SYS_PTRACE; coordinate with security in containerized envs.

Container & Kubernetes Specifics

Map Container to Host PID

docker inspect --format '{{.State.Pid}}' order-service

Then use host PID with top -H, jstack, perf, py-spy.

K8s Pod-Level CPU

kubectl -n production top pod
kubectl -n production top pod order-service-7d9f8c6b4-xk2p9 --containers

Requires metrics-server. Check kubectl -n kube-system get pods -l k8s-app=metrics-server.

Debug Distroless Containers

kubectl -n production debug order-service-7d9f8c6b4-xk2p9 -it --image=busybox --target=order-service

Creates ephemeral debug container with tools; no image rebuild.

CPU Throttling Detection

kubectl -n production exec -it order-service-7d9f8c6b4-xk2p9 -- cat /sys/fs/cgroup/cpu/cpu.stat
# nr_periods 921600
# nr_throttled 45832
# throttled_time 128374920123
nr_throttled/nr_periods ≈ 5%

and growing throttled_time = throttling. Fix: raise limits.cpu from 500m to 2, requests.cpu to 1. Verify nr_throttled stops growing; P99 latency drops from 850ms to 180ms.

Cloud Server Nuances

CPU credits : Burstable instances (AWS T-series) throttle at baseline when credits exhausted — console shows credit balance, not local top.

Shared/entry tiers : Not designed for sustained load; upgrade to compute-optimized.

Cross-AZ latency : Remote DB/cache increases connection hold time → thread pool exhaustion → indirect CPU rise from retries.

Common Misdiagnoses

Symptom : High load average → Wrong Conclusion : CPU bottleneck → Correct Approach : Check %wa and D state processes

Symptom : Process %CPU > 100% → Wrong Conclusion : Process anomaly → Correct Approach : Normal for multi-threaded on multi-core

Symptom : Container top shows host CPU → Wrong Conclusion : Container resource pressure → Correct Approach : Use cpu.stat or kubectl top Symptom : Low CPU but slow response → Wrong Conclusion : CPU not the issue → Correct Approach : Check cloud credits, cgroup throttling, GC pauses, lock waits

Symptom : Single jstack shows thread in method → Wrong Conclusion : Dead loop → Correct Approach : Need multiple samples; sustained stall is key

Tool Permissions & Security

Restrict jstack, perf, strace, py-spy to designated SRE accounts

Clean temp files ( /tmp) containing stack traces with potential PII

Evaluate ptrace (used by strace / py-spy) impact on namespace isolation in multi-tenant clusters

Automated Snapshot Script

Bash script cpu_snapshot.sh collects: uptime, vmstat, mpstat, top processes, memory/IO overview, and if PID provided, thread dump + jstack. Outputs timestamped tarball. Safe (read-only), robust (handles missing tools), version-controlled in team scripts/ repo.

Additional Case Studies

Case 2: Periodic Cron Spike

Hourly CPU spikes (3-5 min) with timeouts. crontab -l reveals 0 * * * * /app/scripts/sync_inventory.sh. Logs show duration grew from 20s to 4+ min as SKUs grew 20k→300k. pidstat -u 1 300 confirms spike alignment. Root cause: per-row DB queries in loop. Fix: batch 500 rows per SQL, reduce frequency to 6-hourly. Post-fix: 40s runtime, CPU spike 90%→30%, timeouts gone.

Case 3: Container Throttling "False Slowness"

K8s migration → P99 latency up, but kubectl top shows 40% CPU. cpu.stat reveals throttling (5% periods throttled, 128s cumulative). Deployment had limits.cpu=500m; bursty parallel sub-tasks (inventory, promo, messaging) exceeded quota. Fix: limits.cpu=2, requests.cpu=1. Throttling stops, P99 850ms→180ms.

Case 4: DB Connection Pool Exhaustion → Indirect CPU Rise

CPU 30%→75%, latency up, timeouts. Threads show distributed moderate usage (not hotspot). Logs: Could not get JDBC Connection. MySQL Threads_connected at max; HikariCP maximumPoolSize=20 saturated. New report query: full scan + filesort (tens of thousands rows, seconds each). EXPLAIN shows type=ALL, Using filesort. Fix: add index on orders.created_at; consider read-replica routing. Post-fix: query ms→ms, pool 5-8 active, CPU back to 30%.

Case 5: Nginx Worker Misconfiguration

Worker processes 40-70% CPU, static assets slow. nginx -T shows worker_processes 2 on 8-core box + gzip_comp_level 9. Fix: worker_processes auto, gzip_comp_level 5, gzip_min_length 1024, worker_connections 4096, epoll, multi_accept on. Reload: nginx -t && nginx -s reload. Workers now 8, CPU per worker single-digit, static latency normal.

Key Terminology Reference

us

— User CPU Time — App compute sy — System CPU Time — Syscalls/scheduling wa — IO Wait — Disk wait si — Software Interrupt — Network packets LWP — Light Weight Process — Linux thread TID — Thread ID — Thread identifier cs — Context Switch — Switch count/sec GC — Garbage Collection — Auto memory mgmt GIL — Global Interpreter Lock — Python thread limit cgroup — Control Group — Linux resource isolation Throttling — CPU limit enforcement — Container quota exceed Burst — Burst credits — Cloud burstable CPU

Post-Incident Actions

Fill monitoring gaps (e.g., cpu.stat throttling, connection pool metrics)

Add scenario to regular load tests (e.g., promo rule surge)

Codify anti-patterns in code review checklist (inefficient regex, missing indexes, row-by-row processing)

Recalculate capacity with real peak data and growth projections

Transform firefighting into fire prevention.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

cgroupperfstraceJava profilingCPU troubleshootingcloud CPU creditscontainer CPU limitsGo pprofKubernetes resource quotasLinux performance analysismonitoring alertsNginx tuningPython py-spy
MaGe Linux Operations
Written by

MaGe Linux Operations

Founded in 2009, MaGe Education is a top Chinese high‑end IT training brand. Its graduates earn 12K+ RMB salaries, and the school has trained tens of thousands of students. It offers high‑pay courses in Linux cloud operations, Python full‑stack, automation, data analysis, AI, and Go high‑concurrency architecture. Thanks to quality courses and a solid reputation, it has talent partnerships with numerous internet firms.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.