Zombie Memcg Exhausting Node Memory? TencentOS's Cross-Kernel Fix
TencentOS analyzes zombie memcg root causes — page cache, shmem, and swap entries — reviews existing solutions' limitations, and proposes a dual-track approach: upstream patches for new kernels (obj_cgroup binding, swap entry fix) and a kernel module for old kernels that asynchronously reparents dying memcg pages and swap entries to online parents, preserving cache while enabling zero-restart deployment.
1. Zombie Memcg: Technical Essence, Impact, and Detection
1.1 Definition and Characteristics
A zombie memcg (dying memcg) is a memory cgroup that has been deleted in user space via rmdir but remains in the kernel because its reference count (refcount) stays above zero. The memcg enters an "offline" state: new processes cannot join, yet the struct mem_cgroup cannot be destroyed due to external resource references.
1.2 Formation Mechanism
Zombie memcgs arise from kernel resources holding references to the memcg:
File Page Cache Binding : Pages from regular file accesses bind their struct page.mem_cgroup to the memcg. If the page cache is not reclaimed (e.g., still accessed or not yet evicted), the reference persists after memcg deletion.
Shmem Shared Memory Pages : Similar to file pages but without a backing file. Without swap, these pages cannot be freed via drop_caches as long as they are mapped by other processes or not deleted.
Swap Entry References : When pages are swapped out, the kernel records the memcg's private ID in the swap entry so that on swap-in the page can be charged to the original memcg (or its parent). This private ID holds a reference on the memcg; as long as the swap entry exists, the memcg cannot be freed — a frequently overlooked path.
Kernel Object Residue : Slab objects (dentry, inode caches) reference the memcg via a memcg pointer. The community's obj_cgroup mechanism has largely mitigated this path.
Different kernel versions vary in which allocation types hold references. Despite newer kernels addressing some types, page cache binding to zombie memcgs remains the primary cause of leaks in production.
Typical triggers include frequent container create/destroy cycles (each leaving page cache references), shmem/tmpfs usage (pages stay resident until file deletion/truncation), and systemd transient cgroups for cron jobs or user sessions that perform file I/O.
1.3 Quantified System Impact
Memory Leak Scale (Linux 5.4.233, measured with drgn): struct mem_cgroup: 2,688 bytes (from kmalloc-4K slab) struct cgroup: ~968 bytes
Per-CPU stats (vmstats_local, vmstats_percpu) on 256-CPU system: ~384 KB
Total per memcg: ~400 KB
At scale:
10,000 zombie memcgs → ~4 GB
NetEase case: 200,000+ zombie memcgs → >80 GB kernel memory
Extreme production clusters: hundreds of GB consumed, causing node memory exhaustion
This memory is not reclaimable by normal mechanisms; reclaim timing is unpredictable and adds overhead.
Traversal Overhead : Kernel operations (global memory stats, scheduling, monitoring) traverse all memcgs including zombies. Latency grows linearly:
Normal: <1 ms traversal
100,000 zombies: hundreds of ms to seconds
Production case: 200,000+ zombies caused memory.stat query latency to jump from 0.5 ms to 1,200 ms, timing out kubelet node status sync
CPU impact: kworker CPU usage rises from <0.1% to 5–10%; extreme cases saturate a CPU core during slab cache traversal
2. Limitations of Existing Solutions
2.1 Community Scan-and-Reclaim Solutions
Two similar approaches use a low-priority kernel thread ( memcg_zombie_reaper) to periodically scan offline memcgs and forcibly free associated pages ( try_free_pages, with drain if needed). Pseudocode illustrates the loop:
Variables: reaper_on = false; pages_scan_limit = large_number; scan_interval = 5 seconds
Function reap_memcg(memcg, background):
if background and memcg recently offline and has dirty pages: return 0
while memcg has pages:
freed = try_free_pages(memcg, 1)
if freed == 0 and not drained:
drain(memcg); drained = true
else if freed == 0: break
if verbose: log memcg status
return freed_pages_count
Function reap_all(background):
reclaimed = 0
for each memcg:
if background and reclaimed >= pages_scan_limit: break
if memcg is online: continue
reclaimed += reap_memcg(memcg, background)
yield CPU
if background: sleep(scan_interval)
Thread reaper_thread():
set low priority
while not stopped:
if reaper_on: reap_all(true)
else: wait until reaper_on or stop
freeze_point()
Init: start reaper_thread; create sysfs entries to control parameters and trigger reapLimitation : Forcibly discarding page cache forces subsequent disk reads, spiking I/O and latency in I/O-heavy workloads (databases, logging).
A more refined community proposal scans and differentiates page types: reclaim clean file pages, reparent valuable cache pages to the parent memcg. This "scan + reclaim/reparent" hybrid inspired later designs.
2.2 Upstream obj_cgroup Restructuring
Introduced in 2020, obj_cgroup replaces direct memcg references on slab objects with a lightweight (tens of bytes) obj_cgroup object. On memcg deletion, obj_cgroup reparents to the parent's objcg_list, allowing the large struct mem_cgroup to be freed. Muchun Song and Qi Zheng proposed extending this to LRU pages (file, anonymous) so pages bind to obj_cgroup and move with it on memcg deletion, eliminating the need for post-hoc scanning.
Drawbacks :
Performance overhead observed: Redis ~6% regression, multi-container worst-case ~23% regression.
Indirection ( objcg → memcg) adds potential overhead on all memory accounting paths.
Residual objects remain (now obj_cgroup instead of mem_cgroup); future growth is uncertain.
Massive kernel restructuring — cannot be delivered as a live patch; requires kernel upgrade and reboot.
Does not address swap entry private ID references; swapped-out shmem pages still pin the memcg.
Despite these, the patch set (33 patches) has been merged into mainline and backported to Linux 6.6.
3. TencentOS Design Principles
Cache Protection vs. Resource Reclamation : Preserve valid cache via reparenting; let the kernel's existing reclaim logic (global watermarks, LRU activity) decide when to actually free pages, avoiding aggressive module-level reclaim that disrupts global memory management.
Completeness : Cover all reference paths — file pages, shmem, slab objects, swap entries — regardless of swap enablement, memory pressure, or container churn frequency.
Compatibility & Stability : Deliver as an independent kernel module for existing kernels (5.4+), no recompile, no reboot, zero business impact. Fine-grained locking (per-page, per-list) avoids global lock contention; state transitions between uncharge and recharge are atomic to prevent races.
4. Diagnosis: drgn Analysis Tool
TencentOS developed a drgn -based script to scan all kernel pages, filter those belonging to dying memcgs, correlate with page mappings to identify memcg paths and file paths, and output per-zombie-memcg top-K hot cached files. Pseudocode:
Initialize caches and data structures
Get the root of the memory cgroup subsystem
For each page in kernel memory:
If scan limit reached, break
Skip slab pages and pages without mapping
Retrieve memcg associated with the page
If memcg is null or invalid, continue
Check if memcg belongs to memory cgroup root; skip if not
Get memcg state; continue if not "ZOMBIE"
Get memcg path
Get file inode path from page mapping
Accumulate total pages and per-file page counts for this memcg
Print summary:
Total pages scanned, dying pages found, total dying memcg count
For each dying memcg: print total cached pages and top-K cached files with page counts5. Solution Details
5.1 New Kernels (6.6): Upstream Patches + Swap Entry Fix
5.1.1 Memcg Charge Mechanism Recap
Charge ( mem_cgroup_charge) binds a page to a memcg: find process's memcg, check quota, set page's mem_cgroup (or memcg_data), increment memcg refcount, update usage stats. Uncharge ( mem_cgroup_uncharge) reverses: decrement stats, decrement refcount, destroy memcg if refcount reaches zero. Key structures: struct page.mem_cgroup (older), struct page.memcg_data (obj_cgroup kernels), struct mem_cgroup.refcount, struct mem_cgroup.lruvec (per-memcg LRU lists protected by spinlock). Special handling for page split (copy memcg to tail page) and swap (record memcg ID on swap-out, charge to nearest non-dying parent on swap-in).
5.1.2 Backported Patch Phases
Audit and fix all folio_memcg() / folio_lruvec() call sites: add RCU read locks or temporary cgroup references where strict folio-memcg binding isn't required, preventing premature memcg free.
Introduce folio_lruvec_lock() lock-retry pattern: after locking, verify folio-lruvec binding consistency; if reparenting in progress ( lruvec_memcg != folio_memcg), unlock and retry. Entire reparent (folio ownership change + LRU list splice to parent) runs under both child and parent per-node lruvec locks. For MGLRU, hotness data is also transferred.
Switch pages to obj_cgroup association: struct page points to struct obj_cgroup instead of struct mem_cgroup. On cgroup offline, memcg_reparent_objcgs() executes under both lruvec locks: splice child LRU lists to parent, transfer MGLRU hotness, reparent objcgs, then reparent stats.
5.1.3 Swap Entry Private ID Fix
Upstream patches solve file/shmem pages but swap entries still hold memcg private IDs, which pin the memcg. TencentOS moves the private ID reference count from mem_cgroup to obj_cgroup (sharing space with an existing bool field, no extra memory):
struct obj_cgroup {
struct percpu_ref refcnt;
struct mem_cgroup *memcg;
atomic_t nr_charged_bytes;
union {
struct list_head list; /* protected by objcg_lock */
struct rcu_head rcu;
};
bool is_root;
refcount_t id_ref; // memcg ID reference shares 8 bytes with is_root
};Now a memcg ID maps to an obj_cgroup pointer, which then yields the memcg. Swap-in of a page from a dying memcg naturally charges to the nearest non-dying parent, and list-lru objects follow the objcg reparent. This fix has been tested and submitted as an RFC to upstream; expected to merge alongside the backported patches in 6.6.
5.2 Old Kernels: Kernel Module with Async Scan & Reparent
For kernels that cannot be upgraded, a kernel module uses delayed_work to periodically scan dying memcgs. Core logic:
for memcg in iter_mem_cgroups():
if memcg is not dying: continue
isolated = isolate_lru_pages(BATCH_COUNT, memcg.lruvec)
putback = []
dst_memcg = get_first_non_dying(memcg)
for page in isolated:
lock_page(page)
// pages locked by others deferred to next round
if !can_reparent(page):
putback.push(page)
// unbind from dying memcg
mem_cgroup_uncharge(page)
// bind to non-dying parent
mem_cgroup_charge(page, dst_memcg)
unlock_page(page)
putback_pages(memcg.lruvec, putback)
if should_throttle(): throttle()
schedule_next_scan(INTERVAL)Configurable via module parameter (e.g.,
echo 5000 > /sys/module/dying_memcg_cleaner/parameters/interval_ms). Pages eligible for lazyfree are freed directly instead of reparented. After all references are reparented or freed, the memcg is released.
For swap entries, the module scans all swap devices, finds entries pointing to dying memcgs, and atomically swaps the memcg ID to the online parent's ID using swap_cgroup_cmpxchg, updating stats and reference counts accordingly:
for swap_entry in for_each_swap_device_cluster():
if swap_count(swap_entry) == 0: continue
memcgid = swap_cgroup_id(swap_entry)
if memcgid == 0: continue
memcg = mem_cgroup_from_id(memcgid)
if !memcg or mem_cgroup_online(memcg): continue
parent = get_online_parent(memcg)
newid = mem_cgroup_id(parent)
if !parent: error += 1; continue
swap_counter_charge(parent, 1)
old_id = swap_cgroup_cmpxchg(swap_entry, memcgid, newid)
assert(old_id == memcgid)
update_stats(memcg, SWAP, -1)
update_stats(parent, SWAP, 1)
swap_counter_charge(memcg, -1)
mem_cgroup_id_put(memcg, 1)Modules are provided for private kernel 5.4+ versions; future integration with TManager platform will automate authorized deployment.
6. Summary and Outlook
Zombie memcg is a kernel resource management challenge in cloud-native environments. TencentOS delivers a dual-track solution covering all reference paths (file pages, shmem, swap entries) across kernel versions: upstream patches for new kernels (obj_cgroup binding + swap ID migration) and a zero-restart kernel module for existing kernels that asynchronously reparents LRU pages and swap entries to online parents, preserving cache. Future work includes addressing obj_cgroup indirection overhead and exploring percpu restructuring or shmem cgroup for more fundamental fixes.
References: Alibaba Cloud Kernel: alinux: memcg: Provide users the ability to reap zombie memcgs Linux MM mailing list: https://lore.kernel.org/linux-mm/[email protected]/ NetEase Tech Blog: https://www.infoq.cn/article/l3qnkhnyusv9w7tgxqvp LWN.net: Fighting the zombie-memcg invasion Oracle Linux Blog: Detecting and debugging zombie memcg issues
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Tencent Architect
We share insights on storage, computing, networking and explore leading industry technologies together.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
