Linux CPU & Memory Tuning: Prevent Freezes, Direct Reclaim Storms & CFS Throttling
This article details Linux kernel memory and CPU tuning techniques for high-concurrency systems, covering watermark scaling to avoid direct reclaim storms, dirty page writeback thresholds, swap and overcommit settings, transparent huge page pitfalls, NUMA zone reclaim, CPU governor and IRQ affinity, and Kubernetes CFS quota throttling solutions.
I. Memory Tuning: The Core Battlefield for Stability and Freeze Prevention
Linux memory management is based on virtual memory, the buddy system, and delayed allocation principles, utilizing all free memory as Page Cache to accelerate I/O. A common misconception is that less free memory means more danger; in Linux, unused memory is wasted memory. The real levers are controlling the rhythm and boundaries of memory reclaim and swap-out.
1. Memory Watermarks and Direct Reclaim Storms
The kernel maintains three watermarks per zone (e.g., ZONE_NORMAL): high (sleep line), low (wake line), and min (direct reclaim floor). These form a state machine:
Safe Zone (Free > High) : Ample memory, allocations instant, kswapd sleeps.
Normal Consumption (Low < Free < High) : Steady consumption, kswapd stays asleep unless previously awakened.
Async Reclaim (Min < Free < Low) : When free memory drops below low , kswapd wakes and asynchronously scans LRU, releasing clean page cache, flushing dirty pages, or swapping cold data — fully in background, no business thread impact .
Direct Reclaim Danger (Free < Min) : If allocation rate outpaces kswapd (e.g., sudden traffic spike creating massive socket buffers), watermark pierces min, forcing Direct Reclaim .
What is Direct Reclaim? Kernel decides physical memory is critically exhausted; background thread cannot catch up. It forcibly suspends the allocating application thread , making it execute synchronous memory reclaim and I/O flush in kernel mode.
During direct reclaim, APIs that normally take milliseconds hang for tens of milliseconds to seconds. In Java/Go this is often misdiagnosed as GC pause, but GC logs show normal times. If reclaim still fails, OOM-Killer triggers.
Key Production Tuning Parameters
vm.watermark_scale_factor (kernel 4.6+) : Controls buffer between low and min (formula: low = min + (managed_pages * watermark_scale_factor) / 10000, default 10 = 0.1%). Tuning : On large-memory servers (64GB/128GB+) or high-burst nodes, increase to 150~200 (1.5%~2.0%), expanding the gap 15-20x so kswapd starts async reclaim much earlier.
vm.min_free_kbytes : Defines the min watermark, reserved for atomic allocations (network interrupts, socket buffers). Tuning : On 64GB-256GB high-concurrency nodes, set to 1GB~4GB (e.g., vm.min_free_kbytes = 2097152). Too small causes packet drops during network bursts; too large wastes usable memory.
2. Dirty Page Writeback: Eliminating I/O Blocking Jitter
Writes are cached in Page Cache as dirty pages, flushed asynchronously by flusher threads. If dirty thresholds are too high, massive dirty pages (tens of GB) accumulate in heavy write workloads (Kafka, Elasticsearch, DB logs). Once the foreground threshold hits, the kernel blocks the application's write call and forces synchronous flush . Limited disk throughput then freezes the entire I/O path for seconds.
Core Parameter Comparison
vm.dirty_background_ratio / vm.dirty_background_bytes : Background async flush threshold (default ~10% of memory). Tuning : Lower for high-throughput writers (e.g., 5% or fixed vm.dirty_background_bytes = 268435456 (256MB)) to make kernel "small steps, smooth flush".
vm.dirty_ratio / vm.dirty_bytes : Forced synchronous flush threshold (default 20%-30%). Tuning : Lower to 10% or fixed 1GB-2GB to avoid massive foreground blocking spikes.
3. Swap Strategy and Overcommit Trade-offs
vm.swappiness True Meaning
Many think swappiness=0 disables swap — completely wrong . Kernel reclaims two page types: file pages (Page Cache) and anonymous pages (heap, stack, private data) . vm.swappiness (0-100, newer kernels up to 200) weights reclaim preference: swappiness=100: Treat file and anonymous pages equally. swappiness=0/1: Strongly avoid swapping anonymous pages, prefer dropping Page Cache. But at memory exhaustion, swap-out or OOM may still occur.
Production advice : For latency-sensitive DBs/middleware (MySQL, Redis, ES), swap-in/out incurs high disk seek and I/O latency (tens of ms). Set vm.swappiness=1 (conservative). On Kubernetes nodes, industry standard is swapoff -a to prevent Kubelet scheduling/eviction misjudgment due to swap.
vm.overcommit_memory : Memory Overcommit and Redis Pitfall
Linux defaults to allowing virtual memory allocations exceeding physical RAM (overcommit):
0 (default, Heuristic) : Kernel tries to satisfy but may refuse if estimated overcommit severe.
1 (Always Overcommit) : Allow any virtual allocation regardless of physical memory. Redis production mandatory : During BGSAVE / BGREWRITEAOF, Redis fork() creates a child process with a full virtual address space clone. Thanks to Copy-on-Write (CoW), actual physical memory increase is tiny, but with overcommit_memory=0 kernel may think full physical memory needed and fail fork() with Cannot allocate memory. Must set sysctl vm.overcommit_memory=1.
2 (Never Overcommit) : Strict limit: virtual allocations cannot exceed Swap + RAM * overcommit_ratio. Suits financial settlement where startup failure is preferred over runtime OOM.
Protecting Critical Infrastructure: oom_score_adj
When OOM-Killer triggers, kernel computes oom_score based on memory usage and penalty weight, killing highest score. Adjusting /proc/<pid>/oom_score_adj (-1000 to 1000) controls survival odds:
Set to -1000 : Process fully immune to OOM-Killer .
In production, kubelet, sshd, core reverse proxies, DB sentinel processes should have OOMScoreAdjust=-1000 via systemd to prevent accidental kill during resource pressure.
4. Huge Pages: Transparent Huge Pages' "Poison" vs Static Huge Pages' "Antidote"
CPU uses page tables for virtual-to-physical translation; TLB (Translation Lookaside Buffer) caches translations. Standard 4KB pages mean tens of millions of entries for tens of GB databases, causing 15%-25% CPU time spent on page table walks. Huge Pages (2MB or 1GB) reduce entries.
Transparent Huge Pages (THP): Database Production Silent Killer
THP merges contiguous 4KB pages into 2MB pages in background. But in Redis, MongoDB, MySQL with frequent alloc/free:
Memory Fragmentation & Compaction Freeze : When contiguous 2MB physical space lacking, khugepaged forces memory compaction, causing global lock contention and allocation stalls — tens of ms to seconds latency spikes.
CoW Memory Amplification : During Redis BGSAVE, if THP enabled, modifying 1 byte forces copying entire 2MB page, exploding physical memory usage tens of times, easily triggering OOM.
Production Iron Rule : On all physical/virtual machines running Redis, MongoDB, Oracle, MySQL, must completely disable THP :
echo never > /sys/kernel/mm/transparent_hugepage/enabled
echo never > /sys/kernel/mm/transparent_hugepage/defragStatic Huge Pages: High-Performance Accelerator
Unlike THP's dynamic merge, Static Huge Pages are explicitly reserved at boot or via vm.nr_hugepages. They are pinned, non-swappable, and eliminate compaction overhead. Use cases : PostgreSQL shared_buffers, DPDK packet processing, high-frequency matching engines. Pinning core caches to 2MB/1GB static huge pages drives TLB hit rate near 100%, slashing CPU address translation overhead and removing fragmentation-induced jitter.
5. NUMA Memory Allocation Trap: zone_reclaim_mode
Modern multi-socket servers use NUMA. Local node memory access is fast; remote node access via QPI/UPI incurs latency/bandwidth penalty. Older kernels sometimes set vm.zone_reclaim_mode=1, meaning: when a NUMA node's local memory exhausts, kernel aggressively reclaims locally (even Direct Reclaim) rather than borrowing free memory from adjacent nodes! This caused classic "MySQL freeze": server has tens of GB free, but MySQL threads on Node 0 trigger local reclaim storms, stalling DB for tens of seconds.
Production Rule : All modern multi-socket servers must ensure vm.zone_reclaim_mode=0 : sysctl vm.zone_reclaim_mode=0 This allows smooth borrowing from other NUMA nodes when local memory low, eliminating stalls from local reclaim deadlocks.
II. CPU Tuning: Full-Throttle Compute, Balanced Interrupts, Protected Cache, Broken Virtualization/Container Limits
If memory tuning focuses on "prevent freeze, prevent swap-out, prevent kill", CPU tuning's core demands are: guarantee compute at full speed, distribute interrupt load evenly, protect CPU caches (L1/L2/L3), and break virtualization/container throttling shackles .
1. Power Management and Frequency Control (Governor & C-States)
Many teams buy 3.0GHz+ servers but find latency misses targets because system runs in default power-saving mode. Modern CPUs have complex power management:
CPU Frequency Governor : OS defaults to powersave or ondemand (dynamic boost). On sudden traffic surge, CPU ramp from hundreds of MHz to turbo takes detection, decision, and hardware lock latency (hundreds of µs to ms), directly causing first-packet long-tail latency.
CPU Sleep Depth (C-States) : Idle cores enter C1, C2, or deep C6 (clock off, voltage reduced). Waking from deep sleep and restoring context costs significant latency.
Production Tuning Practice
For ultra-latency-sensitive gateways, trading systems, core DBs, switch to performance mode, locking CPU frequency, eliminating dynamic scaling jitter:
# Check current governor
cpupower frequency-info
# Force all CPU cores into performance mode
cpupower frequency-set -g performanceIn extreme low-latency (e.g., HFT), add kernel boot params idle=poll or intel_idle.max_cstate=0 to disable sleep, using busy polling for microsecond interrupt response.
2. Interrupt Affinity and Softirq Distribution (IRQ Affinity & RPS/RFS)
Network packets trigger hardware interrupts. By default, all interrupts may be routed to CPU 0 by hardware/BIOS.
Typical Failure Symptom
Overall CPU utilization shows 5%, but per-core view reveals CPU 0 softirq ( si , ksoftirqd) at 100% , packet queues back up and drop ( RX drop), while dozens of other cores idle.
Countermeasures
Physical Multi-Queue NIC Hard Binding : Modern 10G/100G NICs support RSS. High-concurrency nodes usually disable irqbalance service , manually binding each NIC queue's interrupt ( /proc/irq/<irq_num>/smp_affinity) one-to-one to fixed physical cores for hardware-level load balancing.
Cloud VM Virtual NIC Soft Distribution (RPS/RFS) :
RPS (Receive Packet Steering) : Hash source IP/port to distribute softirq across multiple CPU cores.
RFS (Receive Flow Steering) : Track application thread's CPU for a socket, steer softirq to same CPU for maximal cache locality.
In cloud (KVM/Virtio), virtual NICs often have single or few queues. Use kernel soft distribution:
3. NUMA Affinity and Core Isolation (CPU Pinning)
CFS scheduler aims for system-wide "fairness", migrating processes across cores. But on multi-core/NUMA this causes severe degradation:
L1/L2 Cache Instantly Cooled : Thread migration invalidates pre-warmed hot data, causing massive cache misses.
Cross-Socket Memory Haul : Thread scheduled on another NUMA node must access original node's memory over bandwidth-limited interconnect (Intel UPI/AMD Infinity Fabric), latency spikes 2-3x.
Affinity Binding in Practice
Deploying MySQL, Redis single-threaded instances, or HPC nodes, use numactl or taskset to bind process to specific NUMA node and CPU set:
# Start Redis instance, bind to NUMA Node 0, use only Node 0 physical memory
numactl --cpunodebind=0 --membind=0 redis-server /etc/redis/6379.confIn ultra-low-latency finance or DPDK, use kernel boot params isolcpus and nohz_full : isolcpus=24-31: Remove these 8 cores from kernel's regular scheduling domain — kernel never schedules normal processes or softirqs there. nohz_full=24-31: Disable periodic timer ticks (tickless) on these cores, dedicated to trading threads running tight loops for deterministic microsecond latency.
4. Kubernetes Container Era's Biggest Pain: CFS Quota Throttling
In Kubernetes, most engineers set Pod resources like:
resources:
requests:
cpu: "1"
limits:
cpu: "2"This standard config is the #1 culprit for occasional severe interface timeouts (P99 spikes) in multi-threaded high-concurrency apps (Java Spring Boot, Go, Node.js worker pools) .
CFS Throttling Micro-Reality
Kubernetes CPU Limit is not a physical cap but implemented via cgroups CFS (Completely Fair Scheduler) quota:
cpu.cfs_period_us : Scheduling period, default 100,000 µs (100ms) .
cpu.cfs_quota_us : Allowed CPU time per period. If limits: 2, value = 2 * 100ms = 200,000 µs (200ms).
Critical: CPU consumption is summed across all concurrent threads!
Example: Java container with 2-core limit (200ms quota per 100ms period). Burst traffic hits, 10 business threads run concurrently:
10 threads consume 20ms wall time → accumulate 10 * 20ms = 200ms CPU quota.
Quota exhausted, 80ms remaining in period .
Kernel immediately forcibly freezes (Throttled Freeze) all threads in that cgroup !
Container frozen until 80ms passes, next period starts, quota reset, threads resume.
Business sees a 5ms request take 85ms+ due to 80ms freeze. Prometheus shows only 30% CPU usage because of sampling/averaging.
Three Solutions to Break Throttling
Solution 1: Remove CPU Limits for Core Low-Latency Services (Request-Only)
Declare only requests.cpu, leave limits.cpu empty.
Removes cpu.cfs_quota_us straitjacket. Container can "borrow" idle host cores during bursts, eliminating CFS throttling; requests.cpu still guarantees scheduling topology and resource baseline. Standard practice at Uber, Meituan, ByteDance for core online RPC services.
Solution 2: Upgrade to Linux Kernel 5.4+
Old kernels (early 4.14/4.19) have severe CFS quota return contention bugs causing false throttling. New kernels fix this and introduce CPU Burst (burst quota buffer).
Solution 3: Enable Kubernetes CPU Manager Static Policy (Exclusive Physical Cores)
Configure Pod as Guaranteed QoS ( requests.cpu == limits.cpu and integer cores).
Kubelet exclusively binds Pod to specific physical cores, removing it from shared CFS pool — zero throttling plus full L1/L2 cache isolation and NUMA locality.
III. Four Core Scenario Tuning Reference Specs
Tuning without concrete workloads is meaningless. Below summarizes core tuning parameters and underlying logic for four common high-concurrency workloads: Redis, Kafka, MySQL, Kubernetes Node.
1. Redis / Ultra-Fast In-Memory Cache
vm.overcommit_memory = 1: Allow unlimited virtual memory overcommit, ensuring background BGSAVE and AOF rewrite fork() succeed. transparent_hugepage = never: Force disable THP, cure background compaction-induced ms-level freezes and CoW memory explosion. vm.swappiness = 1: Extremely conservative swap, prefer dropping file cache, prevent hot data hitting disk causing I/O stalls.
2. Kafka / High-Throughput Write & Message Queue
vm.dirty_background_ratio = 5(or fixed bytes): Wake background flusher early, continuously and smoothly flush to disk, prevent massive dirty page buildup. vm.dirty_ratio = 10: Lower foreground synchronous write blocking ceiling, avoid burst traffic forcing app threads into flush army causing seconds of unavailability. vm.min_free_kbytes = 2GB+: For 10GbE environments, reserve ample kernel reserved memory to prevent burst consumption or massive packet storms piercing min watermark.
3. MySQL / OLTP Relational Database
vm.zone_reclaim_mode = 0: When NUMA local node memory low, prefer borrowing from other nodes, strictly avoid local reclaim death spiral and Direct Reclaim. cpupower frequency-set -g performance: Lock server to performance mode, disable dynamic power-saving downclocking, eliminate CPU core wake latency during idle-to-peak transitions. innodb_numa_interleave = 1: Enable InnoDB buffer pool interleaved allocation, spread memory pages evenly across NUMA nodes, balance interconnect bandwidth load.
4. Kubernetes Containerized Cluster Nodes
swapoff -a: Fully disable swap cluster-wide, ensure Kubelet accurately evaluates Pod resource quotas and eviction watermarks based on physical memory. cfs_quota monitoring & optimization: Core gateways and online RPC services should use "Request-Only" or integer exclusive cores, and monitor container_cpu_cfs_throttled_periods_total metric. vm.max_map_count = 262144+: Increase max memory map areas per process, prevent Elasticsearch, Java containers, massive coroutine runtimes from crashing due to mmap limit.
IV. Summary and Mental Model
Troubleshooting and governing performance bottlenecks is essentially negotiating with the OS resource management laws. Two ultra-simple mental models capture the global view of Linux tuning:
Network & Filesystem Tuning : Focus on " is the pipe open, how thick is the pipe " — solve connection capacity, data path from small to large, bottlenecks throw clear errors, fix by the book.
Memory Tuning : Bottom line is " don't get swapped out, don't get blocked, don't get killed as the bad guy " —
Don't get swapped out: lower swappiness, eliminate disk I/O jitter.
Don't get blocked: raise watermark_scale_factor, lower dirty page thresholds, disable THP, firmly reject Direct Reclaim and compaction-induced second-level freezes.
Don't get killed: configure overcommit properly, give core guardian processes oom_score_adj = -1000 to prevent mis-kill.
CPU Tuning : Bottom line is " compute doesn't downclock, interrupts don't pick favorites, cache doesn't get blown away, time slices don't get stolen " —
Lock performance governor, disable deep sleep.
Bind multi-queue interrupts and NUMA affinity, protect cache locality.
Beware Kubernetes CFS Quota limits, eradicate periodic freezes caused by concurrent accumulation.
Master this underlying mechanics, and when high-concurrency P99 spikes and sudden freezes strike, you'll no longer blindly "restart and add machines", but wield a scalpel to precisely strike the critical nerves deep in the kernel.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Ops Development & AI Practice
DevSecOps engineer sharing experiences and insights on AI, Web3, and Claude code development. Aims to help solve technical challenges, improve development efficiency, and grow through community interaction. Feel free to comment and discuss.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
