Kubernetes DiskPressure Alerts: Mastering Image, Log & Ephemeral Storage Governance
This guide details diagnosing and resolving Kubernetes node DiskPressure alerts by analyzing filesystem signals (nodefs, imagefs, containerfs), eviction mechanisms, image garbage collection, log rotation, ephemeral storage limits, kubelet configuration, monitoring strategies, and recovery procedures with practical commands for containerd/CRI-O environments.
1. Recognize Three Filesystem Signals
Kubernetes reports three filesystem identifiers: nodefs (node local ephemeral volumes, Pod logs), imagefs (runtime-reported image filesystem, often includes container writable layers), and containerfs (separate container writable layer filesystem when supported). These can point to the same underlying filesystem; mounting /var/lib/containerd on a separate disk does not automatically mean kubelet statistics and ephemeral storage limits respect that layout. Verify via CRI reports, kubelet stats, and mount relationships. PersistentVolumes (PVCs) are not counted in Pod ephemeral-storage usage, but local PVs, hostPath, or node-level storage components can still consume the same physical disk and indirectly cause DiskPressure.
2. Eviction, GC, and Taints Are Distinct Mechanisms
2.1 Hard Thresholds, Soft Grace Periods, and Pressure Debounce
Common Linux node hard thresholds include memory.available: 100Mi, nodefs.available: 10%, nodefs.inodesFree: 5%, imagefs.available: 15%, imagefs.inodesFree: 5%. Actual values depend on the node; containerfs signals have version-derived rules. When a hard threshold is crossed, kubelet first attempts applicable node-level reclamation; only if that fails does it evict Pods. Hard eviction provides no normal Pod termination grace. Soft eviction uses evictionSoftGracePeriod to determine how long the condition must persist; evictionMaxPodGracePeriod limits the Pod termination grace for soft eviction. evictionPressureTransitionPeriod is the debounce time for exiting pressure state, not the soft threshold grace period. Nodes can report pressure before the soft grace period ends; the control plane uses this to manage the node.kubernetes.io/disk-pressure:NoSchedule taint. New Pod scheduling also depends on tolerations, so it is incorrect to say all new Pods are blocked. evictionMinimumReclaim defaults to no extra reclaim amount; do not treat example values like 1Gi/5Gi as defaults.
2.2 Disk Eviction Does Not Follow QoS Class Ordering
Under disk pressure, eviction order is determined by Pod usage on the relevant filesystem, whether usage exceeds requests, Pod priority, and relative usage. Inode pressure lacks an inode request for comparison. CPU/memory QoS classes (BestEffort, Burstable, Guaranteed) do not directly apply to ephemeral-storage eviction. High priority does not exempt a Pod from disk eviction. Node-pressure eviction differs from manual drain; PodDisruptionBudgets do not stop kubelet's node-pressure eviction. Evicted Pods remain as Failed objects; controllers may create replacements, but bare Pods without controllers will not auto-recover.
2.3 Image Garbage Collection
Kubelet's image GC high/low thresholds commonly default to 85/80. Periodic checks reclaim suitable unused images. Minimum age for new images, container references, runtime pinning, shared layers, CRI errors, and GC scheduling all affect reclamation. Do not assume exceeding 85% immediately yields sufficient space. The default imagefs available threshold is 15%, close to the GC high watermark; single-disk setups may hit the imagefs signal simultaneously. Threshold planning must allow time for pulls, extraction, and burst growth. There is no universal kubelet-image-gc-pin field or /var/lib/kubelet/critical-images pin file interface. Locally unique images should be pushed to a recoverable registry; runtime pin support and protection scope must be confirmed per CRI/runtime version.
3. Forensics: Examine Control Plane, Node, and CRI Simultaneously
Control-plane checks (replace NODE):
kubectl get node "$NODE" -o wide
kubectl describe node "$NODE"
kubectl get pods -A --field-selector "spec.nodeName=$NODE" -o wide
kubectl get events -A --field-selector reason=Evicted --sort-by=.metadata.creationTimestamp
kubectl get pods -A --field-selector "spec.nodeName=$NODE,status.phase=Failed" -o json |
jq -r '.items[] | select(.status.reason=="Evicted") | [.metadata.namespace,.metadata.name,.status.message] | @tsv'Condition messages may not include specific thresholds or filesystem names. Combine Pod messages, Events, and kubelet logs; Events can expire and are not permanent audit records. The message The node was low on resource: ephemeral-storage indicates node resource pressure, not necessarily that the Pod exceeded its own limit.
Read kubelet summary stats (requires node proxy permissions; some managed environments restrict this):
kubectl get --raw "/api/v1/nodes/$NODE/proxy/stats/summary" |
jq '{node_fs:.node.fs,runtime:.node.runtime,pods:[.pods[] | {pod:.podRef,ephemeral:."ephemeral-storage"}]}'Record missing data explicitly; do not assume zero.
Node-level commands (adjust paths/service names per deployment):
export LC_ALL=C
sudo journalctl -u kubelet --since '-2 hours' --no-pager | grep -iE 'evict|pressure|garbage|image|disk|rotate'
findmnt -T /var/lib/kubelet -o TARGET,SOURCE,FSTYPE,OPTIONS
findmnt -T /var/lib/containerd -o TARGET,SOURCE,FSTYPE,OPTIONS
findmnt -T /var/log/pods -o TARGET,SOURCE,FSTYPE,OPTIONS
df -hT /var/lib/kubelet /var/lib/containerd /var/log/pods
df -i /var/lib/kubelet /var/lib/containerd /var/log/pods4. Locate with du, But Do Not Equate du to Kubelet Accounting
sudo du -x -B1 --max-depth=1 /var/lib/containerd | sort -nr | head -15
sudo du -x -B1 --max-depth=1 /var/log | sort -nr | head -15
sudo du -x -B1 --max-depth=2 /var/lib/kubelet/pods | sort -nr | head -15
sudo journalctl --disk-usage
sudo lsof +L1 du -xavoids crossing filesystems, but bind mounts on the same filesystem may still be traversed, and it may miss separately mounted business data. Always check mount relationships first. Containerd's snapshotter, remote/lazy pulls, compression ratios, and shared layers affect directory sizes; do not conclude business is writing logs based solely on overlay/content proportions. Recursive scans generate metadata I/O; directories with millions of small files can be slow. Drill down stepwise within confirmed filesystems and directories, optionally with time limits. Differences between df and du can stem from deleted-but-open files, mount masking, reserved space, and accounting granularity; avoid a universal rule like "15% gap means orphan garbage."
Match Pod ownership by UID precisely:
POD_UID=actual-pod-uid
kubectl get pods -A -o json | jq -r --arg uid "$POD_UID" '
.items[] | select(.metadata.uid==$uid) | [.metadata.namespace,.metadata.name,.spec.nodeName] | @tsv'Absence of an API object does not prove the local directory is deletable: kubelet may not have completed unmount, CSI cleanup, or orphan reclamation. Do not directly delete /var/lib/kubelet/pods/<uid> or recursively remove containerd content/snapshot directories.
5. Image Governance: Verify Endpoints, References, and Recoverability
Identify the CRI endpoint on the node. For containerd:
CRI_ENDPOINT=unix:///run/containerd/containerd.sock
CRI=(sudo crictl --runtime-endpoint "$CRI_ENDPOINT" --image-endpoint "$CRI_ENDPOINT")
"${CRI[@]}" info
"${CRI[@]}" images
"${CRI[@]}" images -o json
"${CRI[@]}" ps -a
"${CRI[@]}" ps -a --state exited
"${CRI[@]}" imagefsinfoFirst run crictl --help locally to verify command support. crictl images --size is not a universal exclusive-space accounting method; ctr images ls -o json is not a cross-version copy-paste command. CRI image sizes cannot be simply summed as reclaimable space; shared layers are only released after the last reference disappears.
When GC fails, check kubelet-to-runtime communication, stats validity, existence of reclaimable images, runtime errors, and actual node free space. FreeDiskSpaceFailed may mean insufficient reclaimable images, not necessarily content store corruption.
Prioritize restoring kubelet's reclamation path. If manual cleanup is required, first save container exit codes, logs, and image manifests, then confirm exact objects:
OLD_CONTAINER_ID=confirmed-deletable-exited-container-id
"${CRI[@]}" inspect "$OLD_CONTAINER_ID"
# Only after scene saved and confirmed no longer needed:
"${CRI[@]}" rm "$OLD_CONTAINER_ID"
OLD_IMAGE=registry.example.com/team/app:confirmed-deletable-version
# Only after references, re-pull capability, and change window confirmed:
"${CRI[@]}" rmi "$OLD_IMAGE"Do not use crictl rmi --prune as a periodic governance script. Prune, pin, and runtime RemoveImage behaviors vary by version; assess image re-pullability, registry bandwidth, and offline nodes. Running containers do not protect all rollback versions; retaining zero-replica ReplicaSets does not create node-level container references.
6. Log Governance: Rotation Is the Goal, Not Hard Disk Quotas
In CRI environments, the runtime writes to kubelet-provided Pod log paths; kubelet manages rotation. Typically /var/log/pods holds actual files, while /var/log/containers contains symlinks; do not state "Pod logs are uniformly symlinked to containerd directory." Default containerLogMaxSize is 10Mi, containerLogMaxFiles is 5. Monitoring and rotation are asynchronous; log bursts can exceed target sizes, so "no container exceeds 50Mi" is not guaranteed. kubectl logs usually reads only the current log file, not full history.
Verify config entry and actual rotation results:
systemctl cat kubelet
ps -ww -C kubelet -o args=
# config.yaml path must match actual --config:
sudo sed -n '/containerLog/p' /var/lib/kubelet/config.yaml
sudo ls -lh /var/log/pods/actual-pod-dir/actual-container-dir/Historical rotation filenames may include timestamps and compression suffixes, not just .1/.2. Before cleanup, precisely select the Pod directory, confirm collectors have read the logs, retention policy allows deletion, and files are no longer active. Never delete current logs or let third-party rotation break the runtime log lifecycle.
Application logs written to container writable layer, emptyDir, or hostPath are not automatically managed by containerLogMaxSize. Prefer reducing log volume, using stdout, or application-level rotation. For shared log volumes, use a sidecar in the same Pod to mount and collect; do not generalize to "hostPath can directly mount another Pod's emptyDir."
Journald configuration example (run in root shell; backup existing config first):
mkdir -p /etc/systemd/journald.conf.d
cat > /etc/systemd/journald.conf.d/90-disk-budget.conf <<'CONF'
[Journal]
SystemMaxUse=2G
RuntimeMaxUse=256M
MaxRetentionSec=30day
CONF
systemctl restart systemd-journald
journalctl --disk-usageJournald defaults usually include filesystem-proportional limits; do not claim defaults are unlimited. SystemMaxUse applies to persistent logs, RuntimeMaxUse to runtime logs; directory and Storage behavior depend on distribution. Vacuum only cleans archived logs; running --rotate --vacuum-size=... rotates then deletes matching historical logs — deletion is irreversible, so save failure evidence first.
Audit log parameters use --audit-log-maxbackup, not maxbackups; the example maxsize of 100 is in MB per kube-apiserver docs. New audit logging also requires a valid audit policy, directory permissions, and static Pod volume/volumeMount; merely adding parameters may not suffice. Backup static Pod manifests outside /etc/kubernetes/manifests.
7. Ephemeral Storage Limits: Scheduling Reservation and Usage-Driven Eviction
Pod local ephemeral-storage mainly includes container writable layers, container logs, and disk-backed emptyDir; image read-only layers, PVCs, and hostPath are not counted in the same Pod usage limit. emptyDir.medium: Memory primarily consumes memory.
Requests are the scheduling basis, not exclusive partitions or hard write quotas; without a request, there is typically no effective scheduling budget, not "unlimited" disk. When only a limit is set, a request may be defaulted per default rules; check the final Pod object. Exceeding the limit or disk-backed emptyDir sizeLimit triggers kubelet detection via stats and eviction; it does not guarantee write rejection at the exact overage byte, nor instant enforcement. Directory scanning may miss deleted-but-open files; filesystem project quota accounting can improve stats where supported, but do not equate it to Kubernetes automatically enforcing hard write quotas. Special mount layout stats support must be validated individually.
Deployment example (teaching config only; replace image and values):
apiVersion: apps/v1
kind: Deployment
metadata:
name: cache-demo
namespace: app
spec:
replicas: 2
selector:
matchLabels:
app: cache-demo
template:
metadata:
labels:
app: cache-demo
spec:
containers:
- name: app
image: registry.example.com/team/app:1.4.2
resources:
requests:
memory: 512Mi
ephemeral-storage: 2Gi
limits:
memory: 1Gi
ephemeral-storage: 4Gi
volumeMounts:
- name: scratch
mountPath: /scratch
volumes:
- name: scratch
emptyDir:
sizeLimit: 1GiUsing a Deployment illustrates replacement Pod creation, not a default pattern for stateful workloads. Request values reflect steady state and scheduling budget; limit values reflect peak, fault growth, and acceptable eviction risk — not a uniform "peak × 1.5." Image registry retention policies, business temp file TTL, and cache eviction must also be implemented.
LimitRange can default ephemeral-storage request/limit for new containers; ResourceQuota can govern namespace total budget; neither is a cluster-wide auto-apply mechanism, and they do not retroactively modify existing Pods.
8. Kubelet Config: Merge Into Existing File, Retain Full Thresholds
The following snippet covers typical Linux nodefs/imagefs layout fields; it is not a complete kubelet config. Merge into existing config and preserve cluster's existing memory/pid policies; containerfs nodes require version-specific confirmation.
imageGCHighThresholdPercent: 80
imageGCLowThresholdPercent: 70
containerLogMaxSize: "10Mi"
containerLogMaxFiles: 5
evictionHard:
memory.available: "100Mi"
nodefs.available: "10%"
nodefs.inodesFree: "5%"
imagefs.available: "15%"
imagefs.inodesFree: "5%"
evictionPressureTransitionPeriod: "5m"
evictionMinimumReclaim:
nodefs.available: "1Gi"
imagefs.available: "2Gi"The 80/70 and reclaim values are example governance settings, not defaults or universal optima. Explicitly listing evictionHard may zero out unlisted signals under some default merge behaviors; versions supporting mergeDefaultEvictionSettings can choose merge behavior per docs, but do not add that field to unsupported older versions. Signal keys are flat strings (e.g., nodefs.available), no dot escaping; do not write nodefs: {available: ...}. Command-line args may override config file values; confirm effective config via systemctl cat, process args, and authorized kubelet configz.
Kubelet has no universal "start once, validate YAML, then exit" flow. Do not start an extra kubelet on production hosts for Ansible validation. Perform YAML/schema validation first, then verify startup and behavior on an isolated test node, then roll out production nodes one by one with recoverable backups.
Restarting kubelet usually does not directly terminate existing containers but temporarily affects status and management; no guarantee of zero impact. Containerd shim, Docker live-restore, etc., alter runtime restart impact; cannot universally assert all containers die or all survive.
9. Monitoring: Distinguish Status Gauges and Event Counters
kube_node_status_condition{condition="DiskPressure",status="true"} == 1This is a kube-state-metrics status gauge. kube_pod_status_reason{reason="Evicted"} is also a status gauge; do not use increase(...[15m]) to fake accurate eviction event counts. For event counts, use event collection pipelines or actual kubelet counters, handling event expiration, deduplication, and restarts.
Space, inode, and growth forecasting must cover actual nodefs/imagefs/containerfs mount points; monitoring only /var misses separate runtime disks. node_filesystem_avail_bytes / size_bytes reflects available ratio; its denominator (affected by reserved blocks) differs from df 's Use% and need not be strictly complementary.
Place filesystem watermarks, DiskPressure, image GC failures, log growth, evicted workloads, and collector backlogs on a single timeline. 80% is not a universal alert line; ensure remaining bytes can cover max image pull/extraction, log bursts, and human response time.
10. Recovery and Rollback
Confirm stable disk and inode headroom, DiskPressure transitions to False, related taints auto-clear, and replacement Pods are Ready before uncordoning a previously manually cordoned node:
NODE=worker-01
kubectl get node "$NODE" -o json | jq '{conditions:.status.conditions,taints:.spec.taints}'
# Only if this node was manually cordoned and validation passed:
kubectl uncordon "$NODE"Do not manually delete the disk-pressure taint to restore scheduling. Nodes that cannot be reliably repaired can be rebuilt per node rebuild runbook; first assess local PVs, hostPath, bare Pods, DaemonSets, PDBs, and spare capacity. kubectl drain --ignore-daemonsets --delete-emptydir-data explicitly allows emptyDir data loss and does not evict all DaemonSets/static Pods, so it does not prove "no workloads remain on this node."
Roll back kubelet config by restoring true backups and restarting node-by-node with validation; rolling back LimitRange does not change limits already written into Pod specs — workload templates must be updated. uncordon only restores scheduling; it does not recover deleted data. Image, log, and emptyDir deletions cannot be restored by config rollback.
Drills should run on isolated nodes: controlled log generation to observe rotation, bounded temp file writes to observe usage limits, then verify replacement Pods and collection. Pulling a tiny busybox image and deleting the Pod does not prove GC high-watermark works; do not directly fill an entire disk with dd on a shared test cluster.
11. Postmortem and Long-Term Governance
Teaching example: a node's runtime disk free space dropped; GC could not reclaim enough images. Forensics showed most images still referenced by containers; recent releases added large images, causing old and new versions to coexist. Adjusting GC thresholds cannot create reclaimable space that does not exist; evaluate image slimming, release cadence, node pool capacity, and migratable workloads.
Long-term capacity budgeting must include actual shared-layer usage, release-period coexistence and extraction peaks, log bursts, container working sets, system services, inodes, and safety margins. Sum of ephemeral-storage requests represents only scheduling budget, not peak actual usage, and cannot alone determine disk sizing.
Every postmortem should record: breached signal and actual value, consumption sources, why reclamation fell short, business impact and rebuild outcome, post-fix growth curve, and implemented capacity/quota/rotation improvements. Sustainable automated recovery beats repeated manual deletion.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
MaGe Linux Operations
Founded in 2009, MaGe Education is a top Chinese high‑end IT training brand. Its graduates earn 12K+ RMB salaries, and the school has trained tens of thousands of students. It offers high‑pay courses in Linux cloud operations, Python full‑stack, automation, data analysis, AI, and Go high‑concurrency architecture. Thanks to quality courses and a solid reputation, it has talent partnerships with numerous internet firms.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
