Operations 31 min read

Diagnosing Server Slowdowns: Pinpointing CPU, Memory, Disk, and Network Bottlenecks

This guide provides a systematic approach to diagnosing Linux server slowdowns by distinguishing CPU, memory, disk, and network bottlenecks through evidence-based metric analysis, command-line tools, and practical examples covering load interpretation, PSI, I/O latency, and safe remediation practices.

Ops Community
Ops Community
Ops Community
Diagnosing Server Slowdowns: Pinpointing CPU, Memory, Disk, and Network Bottlenecks

1. Confirm Impact Scope and Collect Scene

Before diving into metrics, answer three questions: Is only one machine slow or all? Is only one endpoint slow or everything including SSH, database, batch jobs? Does the start time correlate with a release, backup, or traffic change? If multiple machines are affected, check shared database, shared storage, network egress, and common changes.

On the target host, run a baseline collection (set LC_ALL=C for consistent output):

export LC_ALL=C
date -Is
uptime
nproc
free -h
vmstat 1 5
mpstat -P ALL 1 3
iostat -y -x 1 3
sar -n DEV,EDEV 1 3
ss -s
sudo journalctl -k --since '-30 min' --no-pager
mpstat

, iostat, sar come from sysstat; if not installed, rely on existing monitoring and basic tools to avoid installing heavy dependencies during an incident. Note: vmstat first report shows boot-time averages; subsequent reports reflect the sampling interval. iostat -y skips the boot-time report. mpstat with an interval reports per-window values, not boot-time averages. These commands are read-only but still consume CPU, I/O, and output bandwidth; limit scope and duration for recursive du/find, packet captures, thread stacks, and profiling. Synchronize timestamps with timezone across hosts; same timezone does not guarantee clock sync.

2. How to Read Key Metrics

2.1 Load Average Reflects Task Count, Not CPU Percentage

Linux Load Average approximates the number of runnable tasks plus tasks in uninterruptible sleep (D state). The 1, 5, 15 minute values are exponentially smoothed trends, not simple arithmetic averages. D state often relates to I/O but can stem from other kernel waits; do not automatically equate to local disk failure. vmstat r shows runnable tasks; b shows tasks blocked on I/O. Sustained high r with busy CPU supports CPU contention; high b requires correlating device latency, task kernel stacks, and remote storage state. A single snapshot cannot establish causality. When comparing CPU counts, consider CPU affinity, cpuset, and cgroup CPU quota — a host with 64 cores does not mean a container limited to 2 CPUs can use all 64.

2.2 CPU Time Breakdown

The following table guides interpretation of CPU time categories:

user / nice — user-space computation, GC, code hotspots; high usage may be normal business logic.

system — kernel paths, syscalls, memory management; requires further sampling, cannot assign root cause by proportion alone.

irq / softirq — hard and soft interrupts; examine per-core distribution; softirq not limited to networking.

iowait — CPU time accounted while waiting for I/O; not disk utilization; low value does not rule out slow I/O.

steal — virtual CPU did not get run time; indicates hypervisor scheduling impact, not proof of host overcommit.

idle — CPU idle; host idle may still hide single-core hotspots or container quota throttling. iowait accounting has limitations on multi-core systems; combine with storage metrics. High context-switch counts are not independent failure evidence; compare with thread count, business load, CPU usage, and latency.

2.3 Memory: Watch Available, Reclaim, and Limits

MemAvailable

estimates memory available for new applications without swapping; it is not a guarantee of immediately allocatable bytes. Page cache is partially reclaimable, but dirty page writeback, unreclaimable slab, locked memory, NUMA local pressure, and container memory limits alter actual allocation behavior.

Swap usage does not equal an active swap storm. Sustained growth in vmstat si/so, major faults, direct reclaim, and business latency anomalies together support a swap-induced performance conclusion. Minor swapping may not cause noticeable stalls; disabling swap may still trigger reclaim and latency before OOM.

2.4 PSI Time Units Are Seconds

Pressure Stall Information (PSI) exposed via /proc/pressure/ shows stall percentages over 10, 60, 300 second windows ( avg10, avg60, avg300) and cumulative microseconds ( total). some means at least some tasks stalled on the resource; full means all non-idle tasks stalled simultaneously. full values must not exceed some on the same window.

for resource in cpu memory io; do
  file="/proc/pressure/$resource"
  if [ -r "$file" ]; then
    printf '
[%s]
' "$resource"
    cat "$file"
  fi
done

Example output:

some avg10=12.50 avg60=8.30 avg300=4.20 total=98372103
full avg10=3.50 avg60=1.20 avg300=0.40 total=18372103

PSI introduced in Linux 4.20; availability depends on kernel config. Newer kernels may show system-level CPU full but it is undefined at system level and usually zero — cannot be used to judge whole-machine CPU stall. I/O PSI high indicates I/O pressure but cannot pinpoint a specific local disk.

2.5 Disk: Latency, Queue, and Throughput

iostat
r_await

/ w_await are average completion times for read/write requests including queue and service time. aqu-sz is average outstanding requests, not pure wait queue length. These are block-device metrics, not equal to database transaction or application request latency. %util on traditional serial devices helps gauge busyness; on NVMe, RAID, and parallel storage it cannot alone determine capacity limits. Queue depth > 0 is normal for parallel loads. Judge bottlenecks by: latency deviation from baseline under same load, sustained queue buildup, throughput hitting device limits, and business impact. Legacy svctm is unsuitable for inferring true service time. Do not mandate a universal "normal latency 1ms" — media, request size, queue depth, sync write semantics, and cloud disk quotas shift baselines. Average latency normality does not rule out tail latency or stuck in-flight requests.

3. CPU Direction: From Process to Thread

First distinguish single-core hotspot, whole-machine compute load, and kernel overhead.

mpstat -P ALL 1 5
pidstat -u 1 5
pidstat -w 1 5
ps -eo pid,pcpu,etime,comm --sort=-pcpu | head -15
ps %CPU

is process lifetime average, not interval sampling; pidstat is not a general syscall counter.

Java thread example (replace with real PID/TID):

APP_PID=12345
HOT_TID=12346
top -H -p "$APP_PID"
TID_HEX=$(printf '%x' "$HOT_TID")
jstack "$APP_PID" | grep -F -A 20 "nid=0x${TID_HEX} "

Must run in correct PID namespace with a compatible, authorized JDK. Thread stacks may introduce pauses; assess impact first; use multiple short sampling to confirm hotspots. Go can sample via authorized pprof endpoints; perf requires controlling frequency, scope, and duration. Do not treat a profiler's fixed overhead percentage as a universal guarantee.

CPU frequency investigation can read cpufreq and driver state, but VM /proc/cpuinfo MHz may not reflect effective compute capacity. powersave governor may not lock to lowest frequency; performance may not guarantee highest. Verify driver, supported policies, thermal limits, and power constraints before changing governor; no cross-distro universal cpufreq service enable command provided.

4. Memory Direction: Separate Usage, Reclaim, and OOM

free -h
vmstat 1 10
pidstat -r 1 5
grep -E 'MemAvailable|Dirty|Writeback|Shmem|SReclaimable' /proc/meminfo
grep -E 'pgscan_direct|pgsteal_direct|allocstall|oom_kill' /proc/vmstat
ps -eo pid,user,rss,vsz,etime,comm --sort=-rss | head -15
sudo journalctl -k --since '-2 hours' --no-pager | grep -iE 'oom|killed process'
/proc/vmstat

fields are cumulative counters; read at intervals to see deltas. nr_dirty / nr_writeback are instantaneous values; note exact field semantics. pidstat -r major faults may come from file-backed mmap reads, not only swap. RSS growth can be cache or working set expansion; only with lifecycle and allocation evidence can a leak be confirmed.

To view per-process swap usage (output in KiB, PID, command):

for proc_dir in /proc/[0-9]*; do
  swap_kib=$(awk '/^VmSwap:/{print $2}' "$proc_dir/status" 2>/dev/null)
  if [[ "$swap_kib" =~ ^[0-9]+$ ]] && (( swap_kib > 0 )); then
    printf '%s\t%s\t%s
' "$swap_kib" "${proc_dir##*/}" "$(cat "$proc_dir/comm" 2>/dev/null)"
  fi
done | sort -rn | head -10
VmSwap

does not cover all shared memory swap-out. Large swap for a process only indicates part of its anonymous memory was swapped; does not prove it currently slows the whole machine.

OOM must be distinguished: global, NUMA/policy limits, cgroup limits, userspace OOM managers. Absence of kernel logs does not rule out OOM — permissions, log rotation, container view, and log forwarding can cause gaps. For cgroup v2 check memory.events, memory.current, memory.max; v1 uses corresponding stat interfaces.

5. Disk Direction: Device, Filesystem, and Waiting Tasks

iostat -y -x 1 5
lsblk -o NAME,TYPE,SIZE,FSTYPE,MOUNTPOINTS
findmnt -T /var/lib/mysql -o TARGET,SOURCE,FSTYPE,OPTIONS
df -hT /var/lib/mysql
df -i /var/lib/mysql
pidstat -d 1 5
ps -eLo pid,tid,stat,wchan:32,comm | awk '$3 ~ /^D/'

Older lsblk may use MOUNTPOINT. pidstat -d process I/O accounting does not always map 1:1 to device throughput due to cache, writeback threads, shared I/O, and network filesystems.

If a task stays in D state, privileged users can read its kernel stack:

APP_PID=12345
TASK_TID=12346
sudo cat "/proc/$APP_PID/task/$TASK_TID/stack"

Stack frames showing NFS/RPC, filesystem, block layer functions are clues, not proof from a single function name. Uninterruptible waits cannot be killed by SIGKILL immediately; some killable waits respond to fatal signals — judge by actual wait path. Prioritize restoring storage path; reboot is not the only answer.

When df shows high usage but du low: check deleted-but-open files, masked data under mount points, filesystem metadata and reserved space, and whether stats cross filesystems. Use sudo lsof +L1 to find deleted open files.

ext4 error handling does not apply to XFS or all filesystems; verify read-only via findmnt mount options and kernel logs. Never run repair tools on mounted production filesystems.

6. Network Direction: Bandwidth, Retransmits, Queues, and DNS

IFACE=eth0  # replace with actual NIC, first run ip -br link
sar -n DEV,EDEV 1 5
ip -s link show dev "$IFACE"
sudo ethtool "$IFACE"
sudo ethtool -S "$IFACE"
ss -lnt
nstat -az TcpRetransSegs TcpOutSegs TcpExtListenOverflows TcpExtListenDrops

Interface drops/errors and TCP retransmits are cumulative counters; examine deltas over the same window. RX dropped may originate from driver, filters, or resource exhaustion — not solely small NIC ring. Retransmits can stem from congestion, reordering, or peer behavior; correlate both ends. ss -lnt listen queue: Recv-Q = completed handshakes awaiting accept; Send-Q = backlog limit. 3/511 not near full; 510/511 warrants attention, confirm with ListenOverflows delta.

sysstat network kB/s definition varies by version; convert to link bit/s carefully (bytes vs bits, KiB vs decimal Mbps). Cloud bandwidth quota may not equal NIC negotiated speed.

cat /proc/net/softnet_stat
sudo sysctl net.netfilter.nf_conntrack_count net.netfilter.nf_conntrack_max
dig +stats example.com
softnet_stat

uses hex cumulative values; second column is dropped, third is time_squeeze (processing budget exhausted, not packet drops). Tuning queue count, ring, RPS, or interrupt affinity may briefly affect link; do not label "hot apply so risk-free".

conntrack module may not be loaded; sysctl entries may not exist. When table fills, inspect entry states and sources; properly closed connections do not all linger as ESTABLISHED with long timeouts. Adjusting TIME_WAIT timeout cannot explain or resolve all ESTABLISHED buildup.

External timing example:

curl -sS -o /dev/null --connect-timeout 5 --max-time 30 \
  -w 'http=%{http_code} dns=%{time_namelookup} connect=%{time_connect} tls=%{time_appconnect} first_byte=%{time_starttransfer} total=%{time_total}
' \
  https://api.example.com/health

These times are cumulative from request start. To isolate stages, subtract; first-byte time includes DNS, connect, TLS, upload, and server wait — not pure application processing. Mid-path mtr loss may be probe rate limiting; last-hop loss needs correlation with probe method and real traffic. Single-point pcap FIN/RST cannot rule out proxy-generated or middlebox-injected packets.

7. Monitoring and Config: Avoid Baking Errors into Scripts

7.1 Use node_exporter Native Counters for Average I/O Time

node_exporter's diskstats collector already provides read/write completion counts and time counters; avoid parsing fixed-column iostat. Example PromQL for average read latency (seconds):

# Average read request time, seconds; NaN when no reads, not faked to zero
rate(node_disk_read_time_seconds_total{device="sdb"}[5m])
/
rate(node_disk_reads_completed_total{device="sdb"}[5m])

Write latency uses

node_disk_write_time_seconds_total / node_disk_writes_completed_total

rate ratio; multiply by 1000 for milliseconds. node_disk_io_time_seconds_total rate approximates busy time proportion, cannot replace read/write latency counters.

Load alert aggregation labels must match:

node_load1
/ on(instance, job)
count by(instance, job) (node_cpu_seconds_total{mode="idle"})
> 2

If environment has cluster or other identity labels, retain them in matching and aggregation. Threshold is a starting point; calibrate by machine role and business latency.

7.2 sysstat History Collection Follows Actual Timer/Cron

systemctl list-timers --all 'sysstat*'
systemctl cat sysstat-collect.timer

Distros may use cron, systemd timer, or /etc/default/sysstat to control collection. History files typically under /var/log/sa or /var/log/sysstat; /etc/sysstat/sysstat is not a universal cron file. sar -A 1 1 only proves immediate sampling works; persistent collection requires verifying scheduled task execution, history file growth, and sar -f readability.

Low-frequency sar CPU/I/O rates are often interval averages; memory fields may be point-in-time samples. Short spikes may be diluted or missed; do not use smooth historical curves to dismiss user reports.

7.3 Changing sysctl Must Be Reversible

Below demonstrates procedure; value 10 is not a universal recommendation. After confirming suitability, in root shell:

backup_dir=$(mktemp -d /var/tmp/swappiness-change.XXXXXX)
sysctl -n vm.swappiness > "$backup_dir/original-value"
sysctl -w vm.swappiness=10
printf 'vm.swappiness = 10
' > /etc/sysctl.d/90-example-swappiness.conf
printf 'Original value recorded in %s
' "$backup_dir/original-value"

Assumes example config did not previously exist; if it did, back up its content first. Rollback using actual backup directory:

backup_dir=/var/tmp/swappiness-change.ACTUAL
original_value=$(cat "$backup_dir/original-value")
sysctl -w "vm.swappiness=$original_value"
rm -- /etc/sysctl.d/90-example-swappiness.conf
sysctl vm.swappiness

Deleting config file does not restore kernel runtime value; original value may not be 60. Modern kernels support swappiness 0–200, representing swap vs file page cost tradeoff; consider zram/zswap and actual storage. dirty_ratio base is memory available for dirty pages, not simple physical memory percentage; bytes and ratio configs for same direction are mutually exclusive. Adjust dirty thresholds after checking existing bytes values and recording original config.

8. Logs Must Pair with Parsing Commands

Nginx log_format goes in http block; use JSON to avoid space-delimited column misparsing:

log_format perf_json escape=json
'{"ts":"$time_iso8601","status":$status,"uri":"$uri",'
'"rt":$request_time,"urt":"$upstream_response_time","upstream":"$upstream_addr"}';
access_log /var/log/nginx/perf.json.log perf_json;

Validate config, reload, generate traffic, then query:

jq -c 'select(.rt > 1)' /var/log/nginx/perf.json.log
$request_time

includes client send/receive and upstream wait; $upstream_response_time is not pure app compute time, multiple upstream attempts produce multiple values. Their difference only aids analysis, cannot be directly labeled "Nginx processing time".

9. Example Retrospective: Memory Pressure Drives Disk Busy

Teaching example, not a verified real incident.

Symptom: API latency rises; same window shows MemAvailable dropping, memory PSI rising, swap in/out increasing, system disk busy. Do not immediately blame bad disk. Continue examining application working set, process swap, cgroup limits, allocation records to determine if memory growth caused paging.

If an app shows abnormal growth, preserve necessary scene, drain traffic, restart or rollback per service runbook, then verify memory curve, swap rate, and P99 latency under same business load. Recovery after restart only suggests process state participated; does not alone prove memory leak.

10. Acceptance and Operational Boundaries

Post-fix must cover at least one recurrence cycle; verify business success rate, P95/P99, resource waits, and error deltas. Do not demand "D state forever zero" or "zero swap"; compare against normal load baseline.

Production troubleshooting boundaries: swapoff may cause severe pressure; available > SwapUsed is not a safety guarantee — consider reclaimability, concurrent growth, and memory policy. drop_caches not for routine memory release; clearing cache increases subsequent read I/O.

Deleting actively written logs may not free space; truncate destroys content, non-append writes may create sparse files. Prefer service-supported rotation and log reopen procedures.

Snapshots, dumps, packet captures must limit size and permissions; failures must be recorded as "not collected", not faked as PASS.

Restarts, parameter changes, bulk rollouts have blast radius. Record original values, change single node first, validate, then expand scope.

The goal of performance troubleshooting is to narrow scope with independent evidence, then verify judgments with verifiable actions. CPU, memory, disk, and network often have causal links; no single metric suffices for attribution.

References

sysstat: iostat manual

procps-ng: vmstat manual

Linux Kernel: PSI

Linux Kernel: softnet_stat source

node_exporter: Linux diskstats collector

Linux Kernel: VM sysctl

Nginx: log module

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

networkLinuxCPUmemoryperformance troubleshootingsysstatPSIdisknode_exporterbottleneck analysis
Ops Community
Written by

Ops Community

A leading IT operations community where professionals share and grow together.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.