Linux Performance Analysis Toolkit: From vmstat to eBPF Tools Explained
This article provides a comprehensive guide to Linux performance analysis tools, covering vmstat, iostat, dstat, iotop, pidstat, top, htop, mpstat, netstat, ps, strace, uptime, lsof, perf, and advanced tools like eBPF, bcc, and Flame Graphs, with usage examples and output interpretation.
Background
The article stems from interest in Linux internals and serves as a benchmark for fundamental computer systems, networking, and OS knowledge. It compiles insights from Brendan Gregg's updated Linux performance tuning tools blog post, explaining principles and tools for system performance optimization.
Background knowledge such as hardware caches and OS kernels is essential because application behavior intertwines with these layers, affecting performance in unexpected ways — e.g., programs unable to utilize cache effectively or excessive system calls causing frequent kernel/user switches.
Performance Analysis Tools
The article references a diagram from Brendan Gregg showing a wide array of tools, each documented via man pages.
vmstat — Virtual Memory Statistics
vmstatmonitors virtual memory, processes, and CPU. Typical usage: vmstat interval times samples every interval seconds for times iterations; omitting times runs indefinitely until stopped with ctrl+c.
The first output line shows averages since boot; subsequent lines show current activity. Key columns:
procs : r = processes waiting for CPU; b = processes in uninterruptible sleep (waiting for I/O).
memory : swpd = blocks swapped to disk; free = unused blocks; buff = buffer blocks; cache = OS cache blocks.
swap : si = blocks swapped in per second; so = blocks swapped out per second.
io : bi = blocks read from block device; bo = blocks written to block device.
system : in = interrupts per second; cs = context switches per second.
cpu : percentages of user ( us), system ( sy), idle ( id), and I/O wait ( wa).
Memory pressure signs: free memory drops sharply, buffer/cache reclamation fails, heavy swap usage ( swpd), frequent paging ( si / so), increased disk I/O ( bi / bo), rising page faults ( in), more context switches ( cs), more processes waiting for I/O ( b), and high CPU I/O wait ( wa).
iostat — CPU and I/O Statistics
iostatreports CPU stats and I/O for devices, adapters, TTY, disks, and CD-ROM. Default output mirrors vmstat CPU info; extended device stats shown with iostat -x. First line = boot averages; subsequent lines = incremental averages per device.
Common metric abbreviations: rq =request, r =read, w =write, qu =queue, sz =size, a =average, tm =time, svc =service. rrqm/s, wrqm/s: merged read/write requests per second (OS merges multiple logical requests into one physical request). r/s, w/s: read/write requests sent to device per second. rsec/s, wsec/s: sectors read/written per second. avgrq-sz: average request size in sectors. avgqu-sz: average queue length. await: average I/O request time (queue + service). svctm: average service time. %util: percentage of time device had at least one active request.
dstat — Versatile System Monitor
dstatdisplays CPU, disk I/O, network, and paging in colorized, readable output, more detailed than vmstat / iostat. Run dstat or with flags like dstat -cdlmnpsy for specific subsystems.
iotop — Per-Process I/O Monitor
iotopshows real-time disk I/O per process, similar UI to top. Non-interactive mode: iotop -bod interval. For per-process I/O details, use pidstat -d interval.
pidstat — Per-Process Resource Monitor
pidstatmonitors CPU, memory, device I/O, task switching, and threads for all or specific processes.
I/O: pidstat -d interval CPU: pidstat -u interval Memory:
pidstat -r intervaltop — System Overview
Summary area shows five categories:
Load: time, logged-in users, 1/5/15-min load averages.
Tasks: running, sleeping, stopped, zombie counts.
CPU: user, system, nice, idle, I/O wait, interrupt percentages.
Memory: total, used, free (system view), buffers, cache.
Swap: total, used, free.
Task area default columns: PID, USER, PR, NI, VIRT, RES, SHR, S, %CPU, %MEM, TIME+, COMMAND.
htop — Interactive Process Viewer
htopis an ncurses-based interactive viewer with color themes, horizontal/vertical scrolling, mouse support, and faster startup than top. Advantages over top: scrollable process list showing full command lines; quicker launch; kill processes without typing PID; mouse interaction.
mpstat — Multi-Processor Statistics
mpstatreads /proc/stat to report per-CPU or average statistics. Common usage: mpstat -P ALL interval times.
netstat — Network Statistics
Displays IP, TCP, UDP, ICMP stats; checks port connections. Common commands: netstat -npl — check if a port is open. netstat -rn — print routing table. netstat -i — interface info: MTU, packets in/out, errors, collisions, output queue length.
ps — Process Snapshot
Many options; refer to man ps. Common patterns: ps aux — list all processes (e.g., for hsserver). ps -ef | grep hundsun — filter processes.
Kill a program:
ps aux | grep mysqld | grep -v grep | awk '{print $2}' | xargs kill -9.
Kill zombie processes: ps -eal | awk '{if ($2 == "Z"){print $4}}' | xargs kill -9.
strace — System Call Tracer
Traces system calls and signals to diagnose abnormal behavior. Example: find which config file mysqld loads: strace -e stat64 mysqld --print-defaults > /dev/null.
uptime — System Uptime and Load
Prints uptime and 1/5/15-minute load averages.
lsof — List Open Files
Lists open files for diagnostics. Common uses: lsof /boot — check filesystem blockage. lsof -i :3306 — find process holding port 3306. lsof -u username — files opened by user. lsof -p 4838 — files opened by PID. lsof -i @192.168.34.128 — remote network connections.
perf — Kernel Performance Profiler
perfis a kernel-integrated profiling tool leveraging tight kernel coupling to access new features early. It identifies hot functions, cache miss rates, and aids optimization. Principle: sampling (e.g., on tick interrupts) captures program context; if a function consumes 90% runtime, ~90% of samples land in its context. High sampling frequency and duration increase reliability.
Common Performance Testing Tools
perf_events : kernel-maintained profiling for apps and kernel.
eBPF tools : use BCC for kernel tracing; eBPF maps managed in user space via bpf syscall.
perf-tools : toolkit built on perf_events and ftrace; few dependencies, supports Linux 3.2+.
bcc (BPF Compiler Collection) : eBPF-based tracing; requires Linux 4.1+ for full features.
ktap : scriptable dynamic kernel tracer, similar to DTrace/SystemTap.
Flame Graphs : visualization for perf, SystemTap, ktap; source at github.com/brendangregg/flamegraph.
Linux Observability Tools
Diagram categorizes tools:
Basic : uptime, top / htop, mpstat, iostat, vmstat, free, ping, nicstat, dstat.
Advanced : sar, netstat, pidstat, strace, tcpdump, blktrace, iotop, slabtop, sysctl, /proc.
Linux Benchmarking Tools
Diagram shows specialized tools for different subsystems (CPU, memory, disk, network, etc.).
Linux Tuning Tools
Diagram focuses on kernel-source-level tuning parameters.
sar — System Activity Reporter
Comprehensive performance reporting: file I/O, system calls, disk I/O, CPU efficiency, memory usage, process activity, IPC. Usage: sar [options] [-A] [-o file] t [n] where t = sampling interval, n = count (default 1), -o file saves binary output, options select subsystems.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Linux Tech Enthusiast
Focused on sharing practical Linux technology content, covering Linux fundamentals, applications, tools, as well as databases, operating systems, network security, and other technical knowledge.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
