Operations 65 min read

Complete Disk I/O Alert Troubleshooting: From Alert to Root Cause with 5 Real Cases

This article details a complete disk I/O alert investigation in production, covering core concepts like IOPS vs throughput, iostat/iotop analysis, and five real-world cases including MySQL missing indexes, log misconfiguration, backup conflicts, Redis persistence, and filesystem mount options, providing a reusable troubleshooting methodology.

Golang Shines
Golang Shines
Golang Shines
Complete Disk I/O Alert Troubleshooting: From Alert to Root Cause with 5 Real Cases

Background

A production database server triggered a disk I/O utilization alert at 95%, causing order system interface timeouts. The article reconstructs the full troubleshooting timeline from alert to root cause, fix, and verification, while extracting a reusable methodology for disk I/O issues.

Core Concepts

IOPS vs Throughput

IOPS measures request count per second; throughput measures data volume (MB/s). They are not linearly correlated: random I/O (database) saturates IOPS first; sequential I/O (logs, backups) saturates throughput first. Optimization direction differs: IOPS bottlenecks need SSD/NVMe or reduced random access; throughput bottlenecks need network storage bandwidth or sequential read/write improvements.

%util Interpretation

iostat

's %util shows device busy time percentage. For HDDs, near 100% means saturation. For SSDs/cloud disks with parallel queues, 100% may not indicate true saturation due to statistical bias. Must cross-check with await, svctm (deprecated in newer sysstat), queue depth.

await and svctm

await

= average I/O request time (queue + service), in ms. Best indicator of "real slowness": HDD baseline ~ms to tens of ms; SSD <1ms. Spikes to tens/hundreds of ms indicate queuing or device slowdown. Focus on await, r_await, w_await to distinguish read vs write latency.

Queue Depth (avgqu-sz / aqu-sz)

Average queued I/O requests. Sustained values >>1 indicate processing capacity lagging request submission. Combine with IOPS and await for severity assessment.

Sequential vs Random I/O

Databases suffer most from random I/O (index scans, small transactions). Sequential I/O (WAL, large files) is disk-friendly. Infer from workload or use blktrace / iostat fine-grained data.

Buffer/Cache Impact

Linux page cache absorbs reads. Memory pressure or O_DIRECT bypass exposes true disk pressure. Check

free -h
available

and buff/cache alongside disk metrics.

Troubleshooting Methodology (Decision Tree)

Confirm alert scope: single/multi host, specific/all devices.

Rule out space exhaustion with df -h (and df -i for inodes).

Assess overall I/O pressure with iostat -x 1 5: watch %util, await, aqu-sz, IOPS, throughput.

Identify culprit process with iotop -o -b -n 5.

Correlate with workload: DB random R/W, log writes, backup jobs, batch processing.

Deep-dive application layer: slow queries, full scans for DB; log flush frequency for apps.

Check storage layer: cloud disk IOPS/throughput quotas; hardware SMART/ dmesg for physical disks.

Confirm root cause, devise emergency stopgap and permanent fix.

Execute fix, verify metrics and business SLAs.

Postmortem: enhance monitoring, update runbooks.

Case 1: MySQL Full Table Scan from Missing Index

Steps

1. Space check: df -h showed 56% and 63% usage — not a space issue.

2. iostat -x 1 5 : Device vdb: %util=98.5%, w_await=68.4ms vs r_await=2.1ms, w/s=892 vs r/s=12, aqu-sz=15.32. Conclusion: write-intensive bottleneck.

3. iotop -o -b -n 5 : mysqld (TID 9201) writing 43.8 MB/s, IO>=92.3%.

4. MySQL analysis: SHOW FULL PROCESSLIST revealed multiple long-running UPDATE order_items SET status=2 WHERE batch_id=... queries (Time >160s). SHOW CREATE TABLE showed no index on batch_id. EXPLAIN UPDATE ... returned type=ALL, rows=2843021 — full table scan.

5. Trigger: Application logs showed a scheduled batch job misconfigured without mutex, causing concurrent executions of the same unindexed update.

Fix

Emergency: KILL the three long queries (warn: rollback generates extra I/O; confirm idempotent retry).

Root: ALTER TABLE order_items ADD INDEX idx_batch_id (batch_id); (Online DDL in MySQL 5.7+/8.0; test in staging, run off-peak). Also fix scheduler mutex.

Verification

EXPLAIN

now shows type=ref, rows=3200. iostat: %util <30%, w_await single-digit ms. iotop no longer shows sustained mysqld write dominance.

Case 2: Debug Log Level Left in Production

Gateway node showed chronic %util 40-60% (3x peers). iotop pointed to gateway process writing ~2 MB/s to access log. Log config ( logback.xml) had root level="DEBUG" with full request/response bodies. Root cause: temporary debug level left after incident. Fix: change to INFO, add rolling policy (size+time, 14 days, 10GB cap, gzip). Verify via watch -n 5 'ls -la /app/logs/gateway-access.log' — write rate dropped to KB/s; %util fell to ~10%. Lesson: track temporary config changes with tickets and reminders.

Case 3: Backup Job Overlapping Business Peak

Daily 02:00 slowdown (10-20 min). crontab revealed mysqldump --single-transaction full backup on primary. iostat during window: %util 20% → 90%+. Root cause: sequential read of entire dataset contended with business I/O. Fix: (1) reschedule off-peak; (2) migrate to replica with replication lag check ( Seconds_Behind_Source <60s). Updated script includes pre-backup lag check, logging, safe cleanup with find -print before -delete. Verification: primary %util stable at 20% during backup window.

Case 4: Redis RDB Snapshot Frequency Mismatch

Cache server: periodic IO spikes ( %util 5% → 90% for seconds). redis-cli CONFIG GET save returned save 900 1 300 10 60 10000. The 60 10000 rule (10k writes in 60s) triggered too often under grown write load. INFO persistence showed rdb_last_save_time updating every 90-120s. Fix: adjust save rules (e.g., 3600 1 300 100 60 10000) or disable if durability not required ( CONFIG SET save ""). Persist via CONFIG REWRITE after backing up redis.conf. Verification: spike frequency matched new rules; client timeouts disappeared.

Case 5: Filesystem Mount Option Difference

New servers 40% slower write throughput than old despite identical hardware. mount diff: old used noatime,nodiratime; new used default relatime. For workload scanning many small files, relatime still incurs metadata writes. Fix: mount -o remount,noatime,nodiratime /data and update /etc/fstab. Re-benchmark: gap closed from 40% to <5%.

Monitoring & Automation

Prometheus Rules

Alert on rate(node_disk_io_time_seconds_total[5m])*100 > 90 for 10m (warning) and write await >50ms for 5m (critical). Thresholds must be tuned per disk type (HDD/SSD/NVMe/cloud).

Healthcheck Script

Bash script ( disk_io_healthcheck.sh) runs df -h, df -i, iostat -x 1 10, flags devices above 80% util (example threshold; production should use historical baselines). Logs to dated file via tee -a.

Ansible Baseline

Playbook checks /data mount options contain noatime via findmnt; fails non-compliant nodes. Dry-run ( --check) and phased rollout ( --limit) recommended for remediation.

Container Considerations

Distinguish container writable layer (storage driver, e.g., overlay2) from mounted volumes; prefer bind mounts/volumes for I/O-heavy data.

Limit container I/O via cgroup: --blkio-weight (relative) or --device-read-bps / --device-write-bps (absolute, at create time). Monitor /sys/fs/cgroup/blkio/ for throttling.

Kubernetes emptyDir without sizeLimit can exhaust node disk/IO; set limits and plan pod density.

Capacity Planning & Documentation

Post-incident: review data growth vs index design, disk spec headroom, backup strategy scalability (incremental, DR), architectural shifts (read replicas, sharding). Use standardized incident template capturing timeline, commands/outputs, root cause evidence, fix links, verification period, rollback log, and follow-up items (monitoring, alerts, code review, capacity review).

Reference Tables

Includes terminology cheat sheet (IOPS, throughput, %util, await, aqu-sz, page cache, O_DIRECT, WAL, redo log, Online DDL, QoS), disk-type baseline ranges (HDD, SATA SSD, NVMe, cloud disks), common error/avoidance table (e.g., don't KILL without verifying source; don't ALTER TABLE at peak; persist Redis config; use logrotate not truncate; phased rollouts; secrets management; multi-metric disk saturation judgment).

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

operationsRedisPrometheusMySQLcapacity planningTroubleshootingansibledisk-ioiotopiostat
Golang Shines
Written by

Golang Shines

We share daily the latest Golang technical articles, practical resources, language news, tutorials, and real-world projects to help everyone learn and improve.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.