Java 27 Compact Object Headers: Real HTTP Benchmark Shows 300 Bytes/Session Saved
This article benchmarks Java 27's default Compact Object Headers against Java 8 and Java 27 with the feature disabled, using identical bytecode. In a Spring Boot authentication service, COH reduces per-session heap by ~300 bytes (10% total heap), with total upgrade savings of ~21%. Benefits scale with active small object count, while hashCode and locking remain fully compatible.
Background
JEP 450 introduced Compact Object Headers (COH) as an experimental feature in Java 24 (disabled by default). JEP 519 made it a product feature in Java 25 (still disabled by default). JEP 534 enabled COH by default in Java 27, so UseCompactObjectHeaders defaults to true.
To answer how much memory COH actually saves, the author conducted two rounds of experiments:
Idealized baseline comparison : Controlled tests with单一 object types (10 million empty objects, 5 million same-length arrays) to measure exact per-object savings and verify runtime mechanisms (hashCode, locking) compatibility.
Real-world scenario validation : Spring Boot authentication/session service and gateway proxy service under real HTTP load to measure overall heap savings.
Both rounds used the same three JVM configurations:
Group A (Baseline) : Corretto 1.8.0_504 (64-bit), no extra flags.
Group B (COH enabled) : Corretto 27.0.0.35.1 (64-bit), default (COH on).
Group C (Control) : Corretto 27.0.0.35.1 with -XX:-UseCompactObjectHeaders.
All test classes were compiled once with javac --release 8 so the same bytecode runs on all three JVMs.
Round 1: Idealized Baseline Comparison
Object Layout Analysis with JOL
Java Object Layout (JOL) prints object memory layout. Key findings for java.lang.Object:
Java 8 (Group A) : mark word 8 bytes + class pointer 4 bytes = 12-byte header. Alignment padding adds 4 bytes → total 16 bytes.
Java 27 COH (Group B) : mark word and class pointer compressed into a single 8-byte header. No padding needed → total 8 bytes.
JOL warns on Java 27 that it detects a "Lilliput VM" and address values are speculative; therefore JOL is used only for structural insight, with cross-verification via heap allocation measurement.
Summary of measured object sizes: java.lang.Object: 16 B → 8 B (saves 8 B: 4 B header + 4 B padding eliminated) byte[1]: 24 B → 16 B (saves 8 B) byte[10]: 32 B → 24 B (saves 8 B) int[10]: 56 B → 56 B (header saves 4 B but padding eats it, total unchanged) String: 24 B → 24 B (same as int[10])
Fixed-Heap Allocation Measurement
Method: Start JVM with -Xms1g -Xmx1g, allocate large numbers of objects (e.g., 10 million), hold references to prevent GC, call System.gc() twice, then read used heap via Runtime.totalMemory() - Runtime.freeMemory(). Subtract the reference array overhead (array header + n × 4 bytes compressed references). Results per object: Object: 16.68 B → 8.45 B (-49%)
POJO (9 mixed primitive fields): 23.46 B → 16.04 B (-32%) byte[10]: 33.15 B → 24.58 B (-26%)
These match JOL structural analysis.
Runtime Compatibility Verification
hashCode : System.identityHashCode() returns stable values before/after COH; object header does not inflate. When the compact header lacks space for the full 31-bit hash, JVM spills hash to a side table, leaving only a marker bit in the header — semantics identical to Java 8.
Locking : Full lock upgrade path tested: synchronized lightweight lock → another thread holds and calls wait() triggering inflation to heavyweight lock → release and re-lock. Throughout, Group B objects stay at 8 bytes; locking, wait, notify semantics work normally.
Round 1 Conclusion
Object header shrinks from 12 to 8 bytes. Zero-field objects halve from 16 to 8 bytes. Small arrays ( byte[1], byte[10]) save 8 bytes each. Objects like int[10] and String see header reduction but total size unchanged due to alignment padding. Large arrays' header savings are negligible. hashCode and locking are fully compatible.
These precise per-object savings form a baseline. Real applications mix many object types (Strings, HashMap.Node, POJOs) where some (like String) don't shrink, so Round 2 tests real services.
Round 2: Real-World Scenario Validation
Experimental Design
Same three JVM groups (A, B, C) run two Spring Boot services built with Spring Boot 2.7.18 (last version supporting Java 8) and maven.compiler.release=8 to ensure identical bytecode. A single fat JAR runs on all three JVMs. Spring Boot 2.7 / Tomcat 9.0.83 runs on Corretto 27 with only native-access warnings at startup, no functional issues.
Both services are designed to have many small objects — the sweet spot for COH benefits.
Service 1: Authentication/Session Service
POST /auth/loginissues token and stores session in ConcurrentHashMap. GET /auth/validate high-frequency validation.
Session object graph mimics real RBAC: ~100 small objects per session including 20 HashMap.Node (permissions), 10 LinkedHashMap.Entry (attributes), ~50 String s, one byte[32], ArrayList and its element array.
Service 2: Gateway Proxy Service
GET /proxy/resourcecollects inbound headers, adds X-Gateway-* / X-Forwarded-*, forwards via RestTemplate to local /upstream/echo (returns ~1 KB JSON + 4 response headers), copies response headers and body back.
Generates many short-lived small objects: LinkedHashMap, HttpHeaders, UUID, byte[]. Upstream response uses DTO EchoResponse.
Load Test & Measurement Method
Independent load generator ( bench.LoadGen, pure JDK HttpURLConnection) runs on JDK 27, shared across all groups.
Auth scenario: 16 threads inject N sessions → 60s random token validation load → call POST /admin/gc to trigger two System.gc() → read heap via Runtime. Resident sessions are old-gen live data; post-GC heap gives precise comparison.
Gateway scenario: 16 threads continuous proxy for 60s, record throughput and latency percentiles, then GC and read heap.
Two heap sizes: -Xms1g -Xmx1g + 200k sessions; -Xms4g -Xmx4g + 500k sessions (32 GB host, 4 GB heap closer to production).
GC logs collected (Java 8: -Xloggc + PrintGCDetails; Java 27: -Xlog:gc).
Measurement endpoint implementation:
@PostMapping("/admin/gc")
public StatsResponse gc() throws InterruptedException {
System.gc();
Thread.sleep(300);
System.gc();
Thread.sleep(800);
return heapStats();
}
static StatsResponse heapStats() {
Runtime rt = Runtime.getRuntime();
return new StatsResponse(SESSIONS.size(),
rt.totalMemory() - rt.freeMemory(), // GC后活跃堆
rt.totalMemory(), rt.maxMemory());
}Two engineering issues encountered during load testing: HttpURLConnection defaults to 5 persistent connections per host ( http.maxConnections); 16 threads exceeded this, causing new connections per request, exhausting Windows ephemeral ports (~16k) with BindException. Fix: set -Dhttp.maxConnections=<threadCount> on both client and server JVMs.
On Windows, -Xlog:gc:file=C:… colon conflicts with unified logging's colon separator; use relative filename with working directory. In PowerShell, $gclog:time is parsed as a scoped variable; write $($gclog):time instead.
Experimental Results
Auth Service: Resident Heap (Post-GC)
Data for 1 GB / 200k sessions and 4 GB / 500k sessions:
Group A (Java 8) : 630.7 MB / 1555.2 MB → 3262 B per session (4 GB run)
Group B (Java 27 + COH) : 502.1 MB / 1232.1 MB → 2584 B per session
Group C (Java 27 - COH) : 560.1 MB / 1375.4 MB → 2884 B per session
Consistent conclusions at both scales:
B vs A (total upgrade benefit) : ~675 B saved per session, ~21% heap reduction.
B vs C (COH net contribution) : ~300 B saved per session (304 B at 1 GB, 300 B at 4 GB), scaling linearly with session count, stable.
C vs A : ~380 B/session additional reduction from other Java 8→27 changes (compact strings, G1 collector, JDK library internals). Therefore, to isolate COH impact, compare B and C — both are JDK 27 differing only in COH switch. The 21% from B vs A is total upgrade benefit, not solely COH.
Breakdown of the ~300 B COH savings per session: Each session has ~100 small objects. The largest contributors are hash table nodes: 20 HashMap.Node + 10 LinkedHashMap.Entry = 30 nodes. In Java 8 each node: 12-byte header + 16-byte fields + 4-byte padding = 32 bytes. With COH: 8-byte header + 16-byte fields = 24 bytes (no padding needed) → 8 bytes saved per node × 30 = 240 bytes. Remaining ~60 bytes come from ~10 other objects (e.g., ArrayList element array) that cross an alignment boundary and shrink by 4-8 bytes each. Sum ≈ 300 bytes, matching measured B-C difference. Round 1 per-object savings accumulate exactly to Round 2 overall savings — the two rounds cross-validate.
GC Behavior (Auth Scenario: Session Injection + 60s Load)
Group A (Java 8, Parallel GC) : 1 GB run: 71 minor GCs, 7.35 s total pause; 4 GB run: 49 minor GCs, 8.25 s total pause.
Group B (Java 27 + COH, G1) : 1 GB: 44 young GCs, 0.88 s total; 4 GB: 14 young GCs, 1.24 s total.
Group C (Java 27 - COH, G1) : 1 GB: 49 young GCs, 0.85 s total; 4 GB: 16 young GCs, 1.56 s total.
Full GC counts in all groups come from startup Metadata GC Threshold and the measurement endpoint's System.gc(); no real full GC during load. Group A's 4 GB run two System.gc() pauses of 3.53 s and 2.66 s dominate its total pause, showing Parallel GC on ~1.6 GB old gen reaches second-level pauses; G1 pauses are markedly shorter. This difference stems from collector generational gap, not COH.
Gateway Service: Throughput & Resident Memory
4 GB run results:
Group A (Java 8) : 5763 rps, p99 14.5 ms, post-GC heap 14.9 MB
Group B (Java 27 + COH) : 12212 rps, p99 5.4 ms, post-GC heap 16.3 MB
Group C (Java 27 - COH) : 7893 rps, p99 10.8 ms, post-GC heap 17.4 MB
Gateway uses mostly short-lived objects; post-GC resident heap is 15-18 MB across all groups — COH impact on resident memory negligible. This aligns with the object-graph model: benefit depends on live small object count , not allocation rate. Young GC counts in 60s: A 34, B 31, C 23 — differences within noise.
Throughput shows high variance: 1 GB run B is 17% lower than C; 4 GB run B is 55% higher than C. Single 60s runs affected by client JIT, connection behavior, etc. Not sufficient for throughput conclusions; core conclusion rests on memory footprint.
Injection Rate (Auxiliary Observation)
Time to inject 500k sessions: A 186 s, B 71.6 s, C 71.4 s. B/C injection rate ~2.6× Java 8, mainly due to JDK version and GC differences (A had 49 minor GCs during injection). COH contribution cannot be isolated from this metric.
Conclusions & Limitations
Conclusions
In a production-like Spring Boot + Tomcat service, COH savings are quantifiable and reproducible: per session (~100 small objects) saves ~300 bytes resident, ~10% heap reduction; total upgrade (Java 8→27) yields ~21%.
Savings proportional to live small object count, not allocation rate. Session caches, config centers, metadata caches with millions of resident small objects are primary beneficiaries; pure forwarding gateways have tiny resident heaps so COH doesn't show in heap usage.
COH has no functional compatibility issues: hashCode side-table, lock inflation, Jackson/reflection all work (verified in Round 1 idealized tests and Round 2 real services).
Limitations
Spring Boot 2.7 + Java 27 is not an officially supported combination (Boot 3 officially supports Java 17+). This test prioritized "same bytecode comparability". Production upgrades with Boot version changes will alter heap profile and object graph.
Throughput/latency are single-run numbers with high variance; not part of core conclusions.
Load generator and service ran on same machine (loopback); absolute throughput not representative of production.
Direct Java 8 vs Java 27 comparison mixes collector, JDK library, framework behavior changes. To attribute to COH alone, compare Group B and C (same JDK 27, only COH toggle).
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
