Operations 117 min read

Linux Disk Space Alert Troubleshooting: Resolve Block, Inode & Deleted File Issues

A comprehensive guide to diagnosing and resolving Linux disk space alerts, covering block vs inode exhaustion, deleted-but-open files, log rotation, filesystem reserved space, read-only recovery, and safe cleanup practices with command examples and monitoring strategies.

Ops Community
Ops Community
Ops Community
Linux Disk Space Alert Troubleshooting: Resolve Block, Inode & Deleted File Issues

Problem Background

Disk alerts are deceptive: unlike CPU alerts, you cannot just scale up. When disk is full, writes fail and services become unavailable. Six common pitfalls:

df vs du mismatch : df -h shows 95% used on / but du -sh /* sums to only 30GB. The missing space is often held by deleted-but-open files.

Space available but cannot write : df -h shows 20% free, yet applications report No space left on device. This is inode exhaustion; use df -i to check.

Deleted files don't free space : Removing a 50GB log leaves df unchanged because the process still holds the file handle. Space is only released when the last holder closes it.

Log growth is the top cause : Business logs without rotation, Nginx access_log not rotated, journald unlimited, Docker container logs unlimited by default.

Filesystem reserved space : ext4 reserves 5% for root. Non-root processes fail at 95% usage while root can still write.

Accidental deletion during troubleshooting : A mistyped rm -rf or wrong find -delete turns a disk alert into data loss.

Core principle: First distinguish block vs inode, then real files vs process-held files, only then act.

Applicable Scenarios

Zabbix/Prometheus disk usage alerts df -h and du -sh discrepancy

Inode exhaustion ( df -i high)

Deleted files not freeing space

Runaway log growth

Docker/Kubernetes node disk alerts

Journald filling disk /tmp or /var/tmp full

Database data directory growing fast

Filesystem mounted read-only

Core Knowledge Points

Block, Inode, Filesystem

Block : Data storage units (default 4KB on ext4). df -h reports block usage. Inode : Metadata (permissions, timestamps, block pointers). Each file consumes at least one inode. df -i reports inode usage. Key : Inode count is fixed at filesystem creation (except XFS, btrfs, tmpfs which allocate dynamically).

Typical inode exhaustion: /var/spool/postfix/maildrop/ with 2 million 1KB files uses only 2GB blocks but consumes all inodes.

Deleted but Open Files

Linux file model: dentry (name→inode), inode (metadata+block pointers), data blocks, reference count (names + open file descriptors). rm removes dentry and decrements link count; if count reaches zero and no process has it open, blocks are freed. If a process still holds the file, blocks remain allocated. ls and du miss it; df still counts it.

Common scenarios: Nginx log deleted but not reloaded, Java app log deleted but JVM holds handle, database temporary files.

Recovery: (1) Restart process (simple, causes downtime). (2) Online truncate via /proc/<pid>/fd/<fd> then signal process to reopen (e.g., nginx -s reopen). Truncation releases blocks but file offset stays; without reopen, next write creates a sparse file and space grows again.

Log Sources

Major log sources and their typical issues:

Business app logs ( /var/log/<app>/, /opt/<app>/logs/): App-controlled, often no rotation or rotation not effective.

Nginx/Apache ( /var/log/nginx/): logrotate config exists but may not trigger or config overwritten.

systemd journal ( /var/log/journal/, /run/log/journal/): Default limit 10% disk; Storage=persistent without SystemMaxUse causes unbounded growth.

Docker container logs ( /var/lib/docker/containers/*/*-json.log): No limit by default ; container writes heavily to stdout.

Database ( /var/lib/mysql/): Binlog via expire_logs_days; binlog not cleaned, slow log too large.

Docker json-file driver has no size limit; must configure max-size and max-file in /etc/docker/daemon.json or per-container --log-opt. Changes only affect new containers.

Filesystem Reserved Space

ext2/3/4 reserve 5% for root. Check with

tune2fs -l /dev/sda1 | grep -i 'reserved block count\|block count'

. Reduce on data disks: tune2fs -m 1 /dev/sdb1. Root filesystem keep at least 3%. XFS has no reserved blocks.

Read-Only Filesystem

Kernel remounts read-only on errors to prevent corruption. Most common cause: disk full causing metadata write failure. Check dmesg -T | grep -iE 'ext4|xfs|read-only|remount|I/O error'. Handling order: (1) Check kernel logs for cause. (2) If space full: free space then remount. (3) If hardware I/O error: backup data, replace disk, do not remount. (4) If metadata corruption: unmount then fsck (ext4) or xfs_repair (XFS). Never run fsck on mounted read-write filesystem.

Overall Troubleshooting Approach

Three-Layer Judgment

Layer 1: block or inode?  df -h high, df -i low → block exhaustion → find large files  df -h low, df -i high → inode exhaustion → find many small files  both high → double exhaustion, handle block first  Layer 2: file really exists or held by process?  du can count → real files → locate directory/file  df vs du gap → deleted-but-open → find process  Layer 3: clean or expand?  Useless data → clean  Valid data → expand (LVM/cloud disk)  Uncontrolled growth → clean + expand + harden

Investigation Chain (8 Steps)

Confirm alert: which host, mount point, usage %, when started.

Resource type: df -h (block), df -i (inode), df -T (fs type).

Mount point ownership: findmnt / lsblk → device, local vs network, LVM vs raw.

Locate directory: du drill down, exclude /proc, /sys, /dev, use -x to avoid crossing mounts.

Locate file: find by size, mtime, type.

Confirm file nature: log? data? temp? process-held? → decide action.

Remediate: clean (carefully), expand, migrate.

Verify: df down, app recovered, no new growth.

Harden: log rotation, journald limits, Docker log limits, alert thresholds.

Time Budget

Suggested time per phase:

Confirm resource type: 1 min

Locate directory: 5 min (use du layer by layer, not full scan)

Locate file: 5 min (use find sort by size)

Confirm file nature: 5 min (check name, lsof, ask app owner)

Stop bleeding: 5 min (truncate known large logs, delete confirmed temp files)

Root cause fix: Varies (configure rotation, expand)

Stop bleeding first : Truncate logs (low risk) before deep analysis. Never run high-risk commands ( rm -rf broad, find -delete without precise conditions) under time pressure.

Practical Steps

Step 0: Capture Alert Info

df -h -x tmpfs -x devtmpfs  # local filesystems only  df -hT  # show fs type  df -i  # inode usage  findmnt -t ext4,xfs -o TARGET,SOURCE,FSTYPE,SIZE,USED,AVAIL,USE%

Interpret thresholds: block >90% critical, inode >85% critical, Avail <5% or <5GB critical. Watch for df hanging on NFS mounts; use df -l for local only.

Step 1: Block vs Inode

TARGET_MOUNT="/"  TARGET_DEV=$(df -h "${TARGET_MOUNT}" | awk 'NR==2{print $1}')  df -h "${TARGET_MOUNT}"  df -i "${TARGET_MOUNT}"  df -T "${TARGET_MOUNT}"  findmnt -o TARGET,SOURCE,FSTYPE,OPTIONS "${TARGET_MOUNT}"

Logic: block high + inode normal → Branch A (large files). inode high + block normal → Branch B (many small files). Both high → Branch A first. Mount options show ro → Branch C (read-only).

Step 2 (Branch A): Locate Large Files

Drill down with du -xhd 1 /mount | sort -hr | head -20 (exclude other mounts). Then

find /mount -xdev -type f -size +500M -printf '%s\t%p
' | sort -rn | head -30

. Use -mmin -60 for recently modified. Exclude known large dirs ( /var/lib/docker, /var/lib/mysql). ncdu -x for interactive browsing (avoid delete key).

Step 3 (Branch A): Confirm File Nature

ls -lh

, stat, lsof -nP "file". Check Modify time: recent → actively written. Identify type by path/name: logs → truncate + reopen; database files → never delete , use SQL ( PURGE BINARY LOGS); temp files → delete if no holder; backups → verify before moving.

Step 4 (Branch A): Handle Process-Held Files

Find deleted-but-open:

lsof -nP +L1 | awk 'NR>1 && $7 ~ /^[0-9]+$/ {print $7, $2, $1, $9}' | sort -rn | head -20

. For each: verify PID and FD target ( readlink /proc/PID/fd/FD), then : > /proc/PID/fd/FD to truncate. Immediately signal reopen: nginx -s reopen, apachectl graceful, systemctl reload rsyslog, or restart app. For Docker:

truncate -s 0 /var/lib/docker/containers/<id>/<id>-json.log

(if file exists) or via /proc if deleted; long-term fix is log limits.

Step 5 (Branch B): Inode Exhaustion

Confirm with df -i. Find inode-heavy directories: du -xhd 1 --inodes / | sort -hr (coreutils 8.22+). Drill down. Classic culprits: /var/spool/postfix/maildrop/ (cron error emails), PHP sessions, /tmp, log rotation leftovers, mail queues, monitoring agent temp files, Docker overlay2 layers, cache dirs, K8s emptyDir. Identify writer via file timestamps, name patterns, content, lsof +D dir. Clean safely: script with dry-run, whitelist, -maxdepth 1, -print0 | xargs -0 rm -f --, double confirmation.

Step 6 (Branch C): Read-Only Recovery

Check kernel logs for cause. If hardware error (SMART Current_Pending_Sector >0, RAID degraded) → backup, replace disk. If space full → free space, then mount -o remount,rw /mount. If metadata corruption → unmount, fsck -n (check only), then interactive fsck or xfs_repair (XFS). fsck -y is dangerous; backup first.

Common Commands

Space Viewing

df -h

– Block usage, human-readable df -i – Inode usage, must run alongside

-h
df -T

– Filesystem type, identify ext4/xfs df -h -x tmpfs -x devtmpfs – Exclude memory fs, real disks only df -h -l – Local only, avoid NFS hang du -xhd 1 dir – First-level subdir sizes, drill down, no cross-fs du -xhd 1 --inodes dir – First-level inode counts, coreutils 8.22+ findmnt – Mount relationships, clearer than

mount
lsblk -f

– Block device tree, LVM/RAID structure stat -f dir – Filesystem stats: block size, inode totals, reserved blocks df -P gives POSIX output (no line wrapping) for scripts. stat -f shows Free vs Available difference = reserved blocks.

File Finding

find dir -xdev -type f -size +500M

– Large files, -xdev no cross-fs find dir -type f -mtime -1 – Modified last day, -mtime days, -mmin minutes find dir -type f -newer ref – Newer than reference, precise comparison find dir -maxdepth 1 – Non-recursive, reduces accidental delete risk find dir -print0 | xargs -0 – Safe for spaces, required for deletion -size default unit is 512-byte blocks; always use explicit suffix ( +100M).

File Handles & Processes

lsof -nP file

– Who opened file, root sees all lsof +L1 – Link count <1 (deleted), key for space recovery lsof +D dir – Processes using dir, slow on large dirs fuser -mv dir – Who uses dir, check before unmount ls -l /proc/PID/fd/ – Process FDs, find deleted links lsof columns: COMMAND, PID, USER, FD (number + r/w/u), TYPE, DEVICE, SIZE/OFF, NODE (inode), NAME (path with (deleted)). For containers: nsenter -t PID -m -u -i -n -p -- lsof +L1.

LVM & Expansion

Order: lvextend then resize2fs (ext4) or xfs_growfs (XFS). Shrink: unmount, e2fsck -f, resize2fs smaller, lvreduce, resize2fs fill. XFS cannot shrink. Never shrink in production ; add disks instead. lvreduce wrong size cuts data irreversibly.

Log Cleaning

Prefer truncate -s 0 file over rm: keeps file, permissions, handles; space released immediately. journalctl --vacuum-size=500M, --vacuum-time=7d, --vacuum-files=5 immediate. docker system prune safe; -a removes unused images; --volumes deletes data volumes (high risk). Always docker system df first to see reclaimable space.

Configuration Examples

logrotate

/var/log/myapp/*.log {  daily  rotate 14  compress  delaycompress  create 0640 myapp myapp  missingok  notifempty  postrotate    /bin/kill -USR1 $(cat /run/myapp.pid 2>/dev/null) 2>/dev/null || true  endscript  maxsize 100M  maxage 30}
create

permissions must match app user. postrotate essential unless copytruncate used (has data loss window, IO overhead, not for huge logs). maxsize triggers rotation mid-day if log spikes. Verify with logrotate -d (dry-run). Ensure /etc/cron.daily/logrotate executable.

journald

[Journal]  Storage=persistent  Compress=yes  SyncIntervalSec=5m  RateLimitIntervalSec=30s  RateLimitBurst=10000  SystemMaxUse=2G  SystemMaxFileSize=200M  SystemMaxFiles=10  RuntimeMaxUse=200M  ForwardToSyslog=no  Seal=no

Default journal limit is 10% of filesystem. Create /var/log/journal for persistent. Check rate limiting: journalctl -b | grep -i 'suppressed\|ratelimit'.

Docker Log Limits

{  "log-driver": "json-file",  "log-opts": {    "max-size": "100m",    "max-file": "5",    "compress": "true"  },  "storage-driver": "overlay2",  "data-root": "/var/lib/docker"}

Only affects new containers; rebuild existing ones. Validate JSON with python3 -m json.tool before restart.

Cleanup Script Template

Safe script features: default dry-run ( --apply to execute), path whitelist, -maxdepth 1, -mtime +7, MAX_DELETE_PER_RUN limit, -print0 | xargs -0, rm -f --, full logging, pre/post usage recording. Run via cron daily in dry-run for weeks before enabling --apply.

Log & Metric Observation

Kernel Logs

dmesg -T | grep -iE 'ext4|xfs|read-only|remount|I/O error|No space'  journalctl -k --since "1 hour ago" | grep -iE 'ext4|xfs|read-only|I/O'

Distinguish info ( recovery complete) from errors. Continuous errors = ongoing issue.

Monitoring Metrics (Prometheus/node_exporter)

node_filesystem_avail_bytes

/ size → block usage (use avail, not free, for non-root) node_filesystem_files_free / files → inode usage node_filesystem_readonly == 1 → immediate alert predict_linear(avail[6h], 4h) < 0 → predicts full in 4 hours (most valuable alert)

Zabbix: vfs.fs.size[/,pfree], vfs.fs.inode[/,pfree], use LLD with vfs.fs.get.

Trend Observation

Hourly cron logging df -h and df -i to file with logrotate. Analyze growth rate: >1%/hour abnormal, stepwise jumps indicate cron jobs.

Investigation Paths

Decision Tree

Start → df -h / df -i / df -T / findmnt  → Mount has ro? → Branch C (read-only)  → Block high? → df vs du gap?  → Gap <5% → Branch A1 (real large files)  → Gap >10% → Branch A2 (deleted-but-open)  → Inode high? → Branch B (many small files)  → Both high → A + B parallel

5-Minute Quick Check Script

Outputs 10 items: block usage, inode usage, mount options, df/du diff, deleted-but-open files, top directories, large files, inode-heavy dirs, kernel errors, recently modified large files. Items 4 and 10 most valuable.

Risk Warnings

High-Risk Operations

rm -rf dir

– Extreme risk, data loss. Verify path, ls first, backup. find ... -delete – High risk, mass accidental delete. Test with -print first. find | xargs rm – High risk, space-splitting filenames. Use -print0 | xargs -0. fsck -y – High risk, may delete data. Run -n first, backup. xfs_repair -L – Extreme risk, log zeroed, metadata loss. Only when explicitly required. lvreduce – Extreme risk, data cut off. Shrink FS first, backup. docker volume prune – High risk, volume data gone. ls + inspect confirm.

/etc/fstab Safety

Use UUID, not /dev/sdX. Add nofail for non-critical mounts, _netdev for network storage. Test with mount -a and findmnt --verify before reboot. If boot fails: GRUB edit → systemd.unit=emergency.target → remount root rw → fix fstab.

rm Safety Rules

Never use undefined variables ( ${DIR:?}), avoid trailing spaces, always cd with || exit, quote variables, use absolute paths, whitelist prefixes, dry-run, double confirmation. Example script includes six validation checks.

Online Truncate Risks

File offset not reset → sparse file, space reappears on next write. Must reopen.

May break apps doing fsync or assuming file position.

Affects concurrent readers (log collectors). Stop collectors first. truncate may fail with permission denied; use shell redirect : > /proc/pid/fd/n as root.

Prefer app-native reopen ( nginx -s reopen) over /proc manipulation.

Cloud-Specific Risks

Expand cloud disk → must grow partition ( growpart), then PV/LV/FS. Step 3 often missed.

Snapshot restore overwrites entire volume; to recover single file, attach snapshot as new disk, mount, copy file.

Never restore snapshot while applications writing.

Verification

Layered Checks

Block: df -h mount – Usage below target

Inode: df -i mount – Usage down

Writable: touch mount/.probe && rm mount/.probe – Success

Mount state: findmnt mount – Shows rw App status: systemctl status svc – Active (running)

App logs: journalctl -u svc --since "5 min ago" | grep -i space – No space errors

App function: curl -w '%{http_code}' url – 200

Kernel: dmesg -T | tail -30 – No new disk errors

Growth trend: Two df 5 min apart – Stable or decreasing

Test writes as app user (not root) because root bypasses reserved blocks.

df/du Consistency

After cleaning deleted-but-open files, df -B1 used and du -xsB1 should converge. Normal diff: few hundred MB to few GB (metadata, reserved blocks). >10% gap warrants investigation.

Continuous Observation

Run watch script for 24h logging usage every 5 min. Stable <1%/day = good. >5%/day = unresolved writer. Stepwise jumps = cron job.

Rollback Plans

Reversible vs Irreversible

Reversible: mount -o remount,ro, tune2fs -m, lvextend (but shrink risky), config files (restore backup + reload). Irreversible: pvresize, resize2fs expand, xfs_growfs, journalctl --vacuum, logrotate -f, truncate, rm, find -delete, docker volume prune, fsck, xfs_repair. Disk cleanups are mostly irreversible → backup, dry-run, whitelist, confirm.

Config Rollback

Always backup before edit ( cp -a file file.$(date +%F-%H%M%S).bak). Restore backup + reload service. For daemon.json: stop docker, restore, start. For journald.conf: restart journald. For logrotate: logrotate -d before -f.

Data Recovery Paths

Backup restore (most reliable).

Cloud snapshot → attach as new disk, mount, copy file.

File recovery tools ( extundelete, photorec) – low success, must unmount immediately.

Database mechanisms: MySQL binlog replay, Redis redis-check-aof --fix.

Replica/rebuild from peers.

Accept loss, rebuild, post-mortem.

Rollback Decision

Rollback if: change causes outage and cause unknown in 5 min, new problem worse than original, data integrity affected. Don't rollback if: fix identified, rollback introduces greater risk (reboot), change mostly complete. Remember: config rollback may not revert runtime state (e.g., restarted containers keep new config).

Production Best Practices

Capacity Planning

Formula: Base OS + (App data growth rate × retention) + (Log growth rate × retention) + Temp peak + 30% buffer. Measure actual growth weekly. Expansion triggers: 70% plan, 80% request, 85% expedite, 90% emergency + clean.

Isolation

Separate mount points: / (50G), /var (100G), /var/log (50G), /var/lib/mysql (500G), /var/lib/docker (500G), /data (2T), /tmp (20G or tmpfs). Prevents log explosion killing database. /tmp as tmpfs:

tmpfs /tmp tmpfs defaults,noatime,nosuid,nodev,size=4G,mode=1777 0 0

(watch memory). Or dedicated partition with systemd-tmpfiles cleanup rules ( d /tmp 1777 root root 10d).

Monitoring Essentials

Block usage (80% warn, 90% critical), inode usage (80/90), readonly (=1 immediate), predict_linear (4h), large file alerts, deleted-but-open total bytes. Collect deleted-open via lsof +L1 every 5-15 min (expensive). Use Prometheus textfile collector with atomic mv.

Maintenance Cadence

Weekly: df block/inode >70%, lsof +L1, top dirs.

Monthly: Reserved blocks check, logrotate working (check .gz files), journal size, docker system df.

Quarterly: Capacity trend analysis, backup restore test (critical), disk SMART health.

Post-Mortem

Timeline, root cause (what, why, why not caught earlier), process review (effective steps, detours, risky actions under pressure), action items (owner, deadline), documentation (runbook, scripts in config mgmt, monitoring updates). Core question: "If this recurs, can we detect and resolve faster?"

Summary

Disk alert troubleshooting hinges on three judgments:

Block or inode? Run df -h and df -i together (5 seconds). Skipping costs 30 minutes.

Real files or process-held? Compare df and du. Gap = deleted-but-open. Locate with lsof +L1 (native filter, better than grep). Truncate via /proc then must reopen (e.g., nginx -s reopen).

Clean or expand? Useless data → clean; valid data → expand; uncontrolled → both + harden.

Key technical preferences: truncate -s 0 > rm (keeps file, perms, rotation config). lsof +L1 > grep deleted (accurate link-count filter). predict_linear alerts > static thresholds (gives hours of lead time). -print0 | xargs -0 mandatory for batch delete (handles spaces). -maxdepth 1 single most effective safety flag. /etc/fstab use UUID + nofail; test with mount -a.

ext4 reserved blocks explain "5% free but write fails".

Never manually delete database files; use PURGE BINARY LOGS etc.

Irreversible ops (cleanup, lvreduce, fsck) require pre-execution risk control: backup, dry-run, whitelist, confirmation.

On-call execution order: 1) df -h/df -i → 2) findmnt/lsblk → 3) df vs du → 4) lsof +L1 → 5) du -xhd 1 → 6) find -size +500M → 7) stat/lsof → 8) truncate logs → 9) /proc/pid/fd + reopen → 10) dmesg → 11) Rotation/expand → 12) Observe stability. Follow sequence, remember: truncate over delete, dry-run over execute, backup over restore.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

MonitoringOperationsLinuxTroubleshootingInodeFilesystemDisk SpaceLog Rotation
Ops Community
Written by

Ops Community

A leading IT operations community where professionals share and grow together.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.