G1 Young Gen Tuning: Why Fixing Size Backfires & How to Adapt
This article explains why fixing G1's young generation size with -Xmn or percentage flags often harms performance, and demonstrates how to tune young GC by analyzing pause targets, Eden sizing, TLAB/PLAB logs, and object promotion patterns to let G1's adaptive sizing balance throughput and latency.
Young GC frequency alone does not indicate a problem. If the system runs high QPS with many short-lived objects, Eden fills quickly, YGC occurs frequently, pauses are short, heap drops noticeably after GC, old generation stays stable, and business P99 latency remains steady — this is normal metabolism.
Real danger signals combine: frequent YGC, increasing single pause times, continuous old generation growth, rising Object Copy cost, and insignificant heap drop after GC. This means temporary objects are surviving YGC and promoting prematurely.
YGC Frequent: First Check Collection Effect
Two YGC patterns look similar in logs but have opposite implications:
"Clean frequently" : Allocation fast, collection fast, old gen stable, business latency smooth. This is just fast metabolism.
"Clean incompletely" : Heap drop not obvious after YGC, Survivor expands, old gen keeps rising, followed by Mixed GC. Young gen is passing pressure to old gen.
Tuning must start from what the business interface can tolerate. Transaction chains and reporting chains need different yardsticks.
Pause Target Too Low: Eden May Shrink Continuously
G1 dynamically adjusts young generation size based on MaxGCPauseMillis. If the target is unreasonably low (e.g., 20ms), G1 reduces Eden regions to meet it. Result: pause shortens but YGC frequency doubles. Smaller Eden → fills faster → more frequent YGC → throughput drops. Worse, short-lived objects may not die before YGC triggers, enter Survivor, cause Survivor bloat → premature promotion → old gen pressure → early Mixed GC → potential Full GC.
Case study: A data reporting system (QPS 1500, P99 80ms) had YGC avg 45ms. Someone lowered MaxGCPauseMillis from 200 to 50. Before: Eden 512M→0, YGC every 8s, old gen stable at 1.2G. After: Eden 256M→0, YGC every 4s, old gen grew from 1.2G to 1.8G, Full GC triggered within a week. Restoring MaxGCPauseMillis to 200 brought Eden back to 512M and stabilized old gen.
Lesson: Don't chase low pause numbers. G1's adaptivity balances pause and throughput, not just minimizes pause. For most online services, 200ms is a reasonable starting point.
Observe young gen shrinkage in logs: [Eden: 128.0M(128.0M)->0.0B(256.0M)] (Eden expanded from 128M to 256M) vs [Eden: 256.0M(256.0M)->0.0B(128.0M)] (Eden shrank from 256M to 128M). If target size (second value in parentheses) keeps decreasing over consecutive GCs, G1 is shrinking young gen to meet pause target. Then check: (1) YGC frequency doubled? (2) Old gen rising? (3) Business latency (P99) worsened? If all three, pause target is too small.
Fixing Young Gen: What Constraints Does G1 Face?
Some dislike G1's adaptivity and fix young gen with -Xmn or -XX:G1NewSizePercent=X -XX:G1MaxNewSizePercent=X. Problems:
1. Traffic Changes, Eden Cannot Adjust
Peak traffic: allocation rate up, Eden fills fast, YGC frequent. Low traffic: Eden same size, YGC frequency drops but single pause lengthens (overhead dominated by scan/prep phases). G1's adaptivity would give Eden more space at peak (if pause budget allows) and shrink at low traffic. Fixed size removes this flexibility.
2. Fixed Young Gen Makes Pause Control Harder
MaxGCPauseMillisworks by estimating historical cost and adjusting CSet size (including Eden regions). Fixed young gen forces G1 to only adjust old region selection. But Young GC must collect entire Eden — that cost is fixed. If fixed young gen too large, single YGC exceeds pause target; if too small, frequent GC and promotion pressure arise. You're fighting G1.
3. Object Lifespan Changes, Fixed Size Becomes Unsuitable
Example: System normally has 90% objects die in Eden. You set fixed young gen accordingly. Later a request-level cache introduces longer-lived objects: only 70% die in Eden, 30% enter Survivor. Fixed young gen unchanged, but GC cost rises (more Object Copy), pauses lengthen. Adaptive G1 would detect increased Object Copy cost and shrink Eden or adjust Survivor ratio. Fixed size leaves you stuck.
Real case: E-commerce order service, 8G heap, -Xmn3G fixed young gen. Normal: YGC 30ms, old gen stable. 618 promotion: traffic 3x, allocation rate spikes. Eden 3G fills fast, YGC from every 10s to every 3s. Single pause still 30ms but frequency too high, throughput drops, P99 from 50ms to 120ms. Removed -Xmn, set -XX:G1NewSizePercent=20 -XX:G1MaxNewSizePercent=60. During promotion, Eden auto-expanded to 4.8G (60%), YGC frequency dropped to every 6s, P99 fell to 60ms. That's adaptive value.
When Is Fixing Young Gen Appropriate?
Suitable only when:
Traffic very stable — no clear peaks/valleys.
Object lifespan uniform — all short-lived or all long-lived, no middle state.
Thoroughly load-tested — confirmed fixed value works under all conditions.
Pause target not strict — not chasing ultra-low pauses.
Typical scenarios: batch jobs, offline computing, internal tools.
Unsuitable for:
Online services — traffic fluctuates.
Complex object lifespans — request-level, session-level, global objects mixed.
Pause-sensitive — strict P99 requirements.
Large heap (>16G) — young gen adjustment space inherently large.
Most online services should not fix young gen.
How to Read TLAB and PLAB Logs
YGC cost includes allocation, not just collection. Each thread has a TLAB (Thread Local Allocation Buffer) in Eden. When TLAB fills, it refills (allocates a new buffer). If Eden too small, TLAB refills frequently. Log example:
[TLAB: gc thread count: 8 slow allocs: 1523 refill waste: 256K max: 512K slow alloc waste: 256K fast refills: 12453]slow allocs — allocations outside TLAB (indicates TLAB insufficient).
refill waste — space wasted during refill.
fast refills — normal refill count.
High slow allocs (thousands) suggests: (1) Eden too small — total space insufficient; (2) Too many threads — each thread gets a slice, total overhead large; (3) TLAB size unreasonable — adjust with -XX:TLABSize.
PLAB (Promotion Local Allocation Buffer) is allocation buffer for Survivor and old gen. Large PLAB allocated indicates high promotion pressure. Combine with Eden size: Small Eden + Large PLAB → objects promote before dying; Large Eden + Small PLAB → promotion normal.
Five-Step Young Gen Troubleshooting
1. Confirm Old Gen Not Rising Continuously
If old gen rises continuously, root cause is object lifespan, not young gen. Tuning young gen only treats symptoms. Find why objects survive YGC: ThreadLocal leaks, cache runaway, request-scoped objects living too long.
2. View YGC Frequency and Single Pause Together
High frequency + short pause + old gen stable → normal, maybe just increase Eden.
High frequency + long pause + old gen rising → many surviving objects, check object lifespan.
Low frequency + long pause + large Eden → possibly high survival rate, check Object Copy.
3. Eden Keeps Shrinking: Check Pause Target First
If Eden continuously shrinks (target size decreasing in logs), check if MaxGCPauseMillis set too low. 200ms is reasonable starting point for most online services. Don't chase 20ms or 50ms unless truly needed.
4. Check Young Gen Bounds Appropriateness
If not fixed, check G1NewSizePercent and G1MaxNewSizePercent. Default 5%~60% gives wide adjustment range. For very large heaps (e.g., 32G), 60% = 19.2G may be too large. Can tighten upper bound: -XX:G1MaxNewSizePercent=40 but leave room for G1 to maneuver.
5. Finally Check TLAB Allocation and Survivor Usage
High slow allocs → consider larger TLAB or Eden. Persistent Survivor bloat → check -XX:SurvivorRatio and object lifespan.
Set Boundaries, Give G1 Adjustment Room
Young gen tuning isn't simply "increase" or "decrease". G1's young gen is dynamic — don't fix it lightly. If you must tune, follow these principles:
First watch old gen — young gen issues often stem from object lifespan.
Don't over-pursue low pauses — 20ms target may shrink Eden, raise frequency.
Don't fix young gen lightly — only if traffic stable, lifespan uniform, thoroughly tested.
Give G1 adjustment room — keep sufficient range between NewSizePercent and MaxNewSizePercent.
Read log details — Eden target size changes, TLAB refill, PLAB promotion.
Frequent YGC isn't necessarily bad; YGC followed by old gen growth is the real danger signal. Tuning goal isn't to eliminate YGC but to make it efficient, stable, and not create pressure on old gen. Don't bind G1; let it adapt. Your job is to set boundaries, then observe its behavior within them. If it can't adapt within boundaries, either boundaries are wrong or object lifespan has issues — that's the root problem to solve.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
