JDK 25 G1GC Bug Steals Pinned Objects: Silent Data Corruption in Production
Ctrip's JDK 25 upgrade revealed a silent data corruption bug in G1GC's Optional Evacuation phase, where pinned objects used by JNI critical sections (e.g., zstd-jni) were incorrectly moved, corrupting Parquet/ORC files without write-time errors; root cause traced to commit 86cec4ea and fixed via backport JDK-8377811.
Background
Ctrip's big data platform runs large-scale Spark and Flink clusters on JDK 21. To leverage JDK 25 LTS features — especially Compact Object Headers (JEP 519, -XX:+UseCompactObjectHeaders) for memory savings and GC efficiency — they adapted Spark for JDK 25 and began canary rollout.
During canary, users reported read failures on Parquet and ORC files written by Spark and Flink. Write paths showed no errors; CRC checks passed. Corruption surfaced only during downstream reads.
Impact Scope
The bug affects all released JDK 25 versions (25.0.0, 25.0.1, 25.0.2), fixed in 25.0.3 (expected 2026-04-21). Any Java application using G1GC (JDK 25 default) is at risk. Workarounds: -XX:+UseParallelGC or -XX:+UseZGC.
Root cause: G1 GC's Optional Evacuation incorrectly moves objects pinned by JNI critical sections ( GetPrimitiveArrayCritical / ReleasePrimitiveArrayCritical). Affected libraries include zstd-jni (used by ORC, Parquet, Kafka), JDK's built-in Zip/Deflate ( java.util.zip.Deflater / Inflater), and any native compression/encryption/math libraries using JNI critical arrays.
Danger: silent corruption — writes succeed, data is persisted corrupted, errors appear only on read/decompress; data is unrecoverable.
Observed Read Errors
Downstream jobs failed on specific columns with Zstd decompression exceptions: Src size is incorrect (from ZstdDecompressCtx.decompressByteArray)
Decompression error: Destination buffer is too small Decompression error: Corrupted block detectedCorrupted File Analysis
4.1 Locating Damaged Columns
Used select sum(hash(struct(colX))) on Parquet/ORC files to pinpoint damaged columns (columnar storage means reading other columns succeeds).
4.2 Ruling Out Storage Media
Parquet page checksums ( parquet.page.write-checksum.enabled=true, parquet.page.verify-checksum.enabled=true) verified via parquet-cli — no errors, confirming bytes on disk match what was written; HDFS not at fault.
4.3 Verifying Corruption Characteristics
Extracted compressed bytes, ran zstd -d → Decoding error (36) : Data corruption detected; crc32 matched page header checksum.
Modified Parquet to skip damaged pages — only a few pages in a few columns corrupted; rest readable.
4.4 Recovery Attempts
Compiled zstd with DEBUGLEVEL=5; AI analysis of debug logs yielded no clear cause.
Cursor AI analyzed raw binary, confirmed Zstd frame structure, generated Python script to recover partial data using Zstd's block independence.
Suspected Directions
JDK 25 Compact Object Headers altering object layout.
zstd-jni, Parquet, ORC, Spark, Flink not fully adapted to JDK 25.
OS/kernel compatibility.
Other environment/component interactions.
Reproduction
ORC Zstd has two implementations: airlift aircompressor (pure Java) and zstd-jni (C, default). Switching to Java impl ( -Dorc.compression.zstd.impl=java) eliminated the issue, implicating zstd-jni.
Physical clusters rarely reproduced; Docker-based clusters (Spark Executors in Docker) reproduced intermittently. OS/kernel differed.
Code-Level Attempts
Enabled Zstd checksum ( zstdCompressCtx.setChecksum(true)) — no effect.
Implemented compress-decompress-verify-recompress loop: compressed lengths matched, but byte offsets and corruption locations were non-deterministic; could not isolate root cause.
JDK-Level Attempts & Bisection
Disabled Compact Object Headers — issue persisted.
All JDK 25 releases (25.0.0, 25.0.1, 25.0.2) reproduced.
JDK 23, JDK 24 (latest) did not reproduce.
JDK 26/27 untested due to Spark incompatibility (removed jdk.internal.ref.Cleaner).
Switched GC: ParallelGC and ZGC on JDK 25 did not reproduce; only G1GC triggered corruption.
Bisection across JDK 25 early-access builds: jdk-25+9 clean, jdk-25+10 reproduced.
Built custom JDKs from specific commits via GitHub Actions workflow (Dockerfile based on CentOS 7 for glibc 2.17 compatibility, embedding commit ID in java -version output).
Identified offending commit: 86cec4ea (JDK-8343782: "G1: Use one G1CardSet instance for multiple old gen regions", resolved in build b10). Previous commit 006ed5c0 clean.
Configuration Test Results
JDK 25 + -XX:+UseG1GC: Data corruption (compressed data corrupted)
JDK 25 + -XX:+UseParallelGC: Works correctly
JDK 25 + -XX:+UseZGC: Works correctly
JDK 21 + -XX:+UseG1GC: Works correctly
JDK 25 (commit 86cec4ea) + G1GC: Data corruption (confirms this commit is root cause)
JDK 25 (commit 006ed5c0) + G1GC: Works correctly
Root Cause Analysis
9.1 G1 GC Collection Phases
Failure occurs in G1 Mixed GC's Optional Evacuations phase (introduced in Java 12, JEP 344). G1 splits Mixed GC into:
Mandatory Part : all young regions + some old regions required for progress — must complete in this pause.
Optional Part : remaining old-region candidates — only evacuated if time budget allows.
9.2 Why JDK-8343782 Causes Corruption
Commit 86cec4ea aimed to reduce G1 memory overhead but introduced a critical bug: during Optional Evacuation, it fails to respect has_pinned_objects() flag on regions.
Impact chain for zstd-jni:
zstd-jni uses GetPrimitiveArrayCritical / ReleasePrimitiveArrayCritical to get raw pointer to Java byte[] for compression.
While pinned, the containing heap region is marked has_pinned_objects.
Pinning tells GC: "Native code is reading/writing this memory address — do not move it."
Post-optimization G1 ignores the pin flag during Optional Evacuation and relocates the array.
Native compression writes to the old address (now stale) or memory becomes scrambled.
No exception thrown (raw memory write). Corrupted bytes persisted to disk. Later decompression fails due to invalid Zstd format.
Summary : pinned flag lost → GC moves pinned object → JNI raw pointer points to freed memory → silent data corruption.
9.3 Why JDK-8370807 Fixes It
Commit JDK-8370807 ("G1: Improve region attribute table method naming", build b22) appears to be a refactor but corrects the missing has_pinned_objects() check in Optional Evacuation registration code. It ensures regions with JNI-pinned objects have the attribute properly set and honored, preventing evacuation of those regions. Object stays put; native read/write addresses stay valid; compression/decompression works.
Retrospective
10.1 Why Some Clusters Didn't Reproduce
Resource-rich YARN clusters: low executor reuse, fewer GCs → hard to hit. Resource-constrained clusters: high reuse, frequent GCs → easier to trigger. Verified by limiting spark.dynamicAllocation.maxExecutors on large cluster to force reuse — reproduced.
10.2 Broader Impact
Not limited to zstd-jni. Any JNI critical-section usage is vulnerable. Verified with JDK's built-in Zip/Deflate (also JNI-based) — reproduced
java.util.zip.DataFormatException: invalid stored block lengthsand Bad compression data errors in ORC ZlibCodec.
AI-Assisted Debugging Recap
Zstd log analysis : Data analysis assistant — extracted key signals from tens of thousands of log lines, narrowed search direction.
Zstd binary analysis : Format parsing expert — parsed Zstd frame structure, generated recovery scripts.
JDK build environment : Code generation accelerator — generated GitHub Actions workflow, fixed glibc compatibility (CentOS 7 base).
JDK commit search : Code repository search engine — reduced candidate commit range for bisection.
Root cause understanding : Technical translator — helped interpret G1 internals, draft bug reports.
Tools used: Cursor, GitHub Copilot, Kiro, Gemini, Claude. Cursor analyzed corrupted binaries and wrote recovery scripts; GitHub Agents produced the JDK build workflow ("vibe coding"), solved cross-environment glibc issues; multiple AIs searched massive JDK repo for G1 commits after "JDK 25 + G1GC" hypothesis formed.
Summary
Initial suspicion never pointed to a JDK GC bug — silent data corruption from GC moving pinned objects was unexpected. Thanks to engineers from NetEase, eBay, Bilibili for collaborative debugging, eliminating red herrings, and helping pinpoint the issue.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
ITPUB
Official ITPUB account sharing technical insights, community news, and exciting events.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
