Apple A20 Pro Die Analysis: 2nm GAA, 36MB SLC, Planar vs 3D Stacking
Kurnal Insights' die shots reveal Apple A20 Pro's 98.72mm² TSMC N2 die with ~260B transistors, 7-core GPU, 36MB unified SLC, 2P+4E CPU, and WMCM packaging, contrasted against Kirin 9050 Pro's 3D-stacked architecture and traced across A18/A19/A20 generational evolution.
Overall Layout Overview
Kurnal Insights' die analysis estimates the Apple A20 Pro at approximately 260 billion transistors (industry estimate, not official) fabricated on TSMC N2 (2nm GAA, second-generation nanosheet) process. The die measures 98.72 mm², nearly identical to A19 Pro, yet achieves cache and core upgrades within the same footprint. The key architectural change is SLC expansion from 24 MB (A19 Pro) to 36 MB, enabled by N2 density gains. The SoC uses WMCM (Wafer-Level Multi-Chip Module) side-by-side packaging, leaving the die backside free for direct vapor-chamber contact, significantly improving sustained thermal performance.
The die adopts a top-bottom partition with unified I/O on the right edge:
Top region: GPU cluster + System Level Cache (SLC)
Middle region: CPU big/little clusters (2P+4E) + CPU L2 caches
Bottom-left: Neural Engine (ANE) + Media Engine
Bottom-right: ISP, video engine, Secure Enclave, peripheral controllers
Far-right strip: LPDDR5X memory PHY (3×32-bit = 96-bit total)
Module-by-Module Breakdown
1. GPU Graphics Unit
7 GPU cores (Core #1–#7), up from 6 in A19 Pro; core count unchanged but microarchitecture optimized for ray-tracing and Mesh Shader performance.
GPU L2 1 MiB : shared L2 cache composed of four 256 KB blocks, centrally located among the 7 cores.
The GPU cluster occupies the largest die area, spanning the entire upper half.
2. SLC (System Level Cache)
Three 12 MiB slices (SLC #1, #2, #3) totaling 36 MB — the single biggest upgrade over A19 Pro's 24 MB.
SLC is a unified cache shared by CPU, GPU, ANE, and ISP, dramatically reducing external memory latency and underpinning Apple's SoC performance advantage.
SLC controllers sit adjacent to each slice.
3. CPU Cluster
Architecture remains 2P+4E (6 cores total) .
P-Cores #1 & #2 : two new high-performance cores. P-Core L2 cache: central 16 MB + 8 MB on each side = 32 MB total (unchanged from A18/A19).
E-Cores #1–#4 : four efficiency cores. E-Core L2 cache: 8 MB (up from 6 MB in A19 Pro), paired with E-AMX vector extension units.
4. ANE (Neural Engine)
Two independent ANE blocks (#1, #2), each with 16 cores = 32 cores total , each backed by 2 MB ANE cache.
Numerous ANE sub-arrays (ANE 1&2, 4&6…) handle on-device LLMs, AI image processing, and NPU inference.
Apple's ANE is a dedicated hardware AI accelerator, not sharing compute with GPU.
5. Multimedia & Imaging
ISP: image signal processor for raw camera data.
Video Engine: hardware video encode/decode.
Media Engine: media processing unit.
Secure Enclave: isolated security processor storing biometric keys and encrypted data.
eFlash: embedded one-time programmable flash for firmware and security keys.
6. I/O & Peripheral Controllers (Bottom + Far Right)
LPDDR5X PHY: 3×32-bit = 96-bit, 115.2 GB/s peak bandwidth.
USB Controller/PHY, Display controller, 3-lane PCIe, 6-lane ASDCI (Apple proprietary high-speed serial).
A20 Pro vs Kirin 9050 Pro: Die Architecture Comparison
Data sourced from Kurnal Insights die teardowns; Kirin 9050 Pro reference: Kirin 9050 Pro Die Shot Deep Analysis (WeChat article).
Die Paradigm : A20 Pro — monolithic 2D planar die (all logic + SRAM on one silicon piece, traditional planar SoC design). Kirin 9050 Pro — dual-die vertical stacking (Logic Folding): compute die + SRAM cache die, Cu-Cu hybrid bonding, two dies stacked vertically, functioning as one SoC.
Process : A20 Pro — TSMC N2 (2nm GAA). Kirin 9050 Pro — domestic DUV process; dual-layer stacking boosts effective transistor density to 238 MTr/mm².
Die Area : A20 Pro — single die: 98.72 mm² . Kirin 9050 Pro — single die: 120.65 mm² × 2 dies.
CPU Architecture : A20 Pro — 2P+4E (6 cores), 2 performance + 4 efficiency, no SMT. Kirin 9050 Pro — four-cluster 9-core/16-thread: 1 super-big + 2 big + 4 mid + 2 little, SMT supported; Lingxi CPU architecture.
CPU Cache : A20 Pro — P-core L2 32 MB, E-core L2 8 MB; unified SLC 36 MB (all compute units share). Kirin 9050 Pro — hierarchical private caches per core cluster (L1/L2/L3); SLC 12 MB ; SRAM die dedicated to large caches.
GPU : A20 Pro — 7-core custom GPU, hardware ray-tracing, integrated at top of single die. Kirin 9050 Pro — Maleoon 955 GPU on compute die upper region; graphics & ray-tracing significantly improved.
AI NPU : A20 Pro — ANE (dual 16-core = 32 cores), independent AI hardware. Kirin 9050 Pro — Da Vinci NPU (4 cores) on compute die, paired with SRAM die private cache.
Memory Interface : A20 Pro — LPDDR5X, 96-bit, right-edge PHY strip, 115.2 GB/s. Kirin 9050 Pro — LPDDR5X, 64-bit; integrated Balong 5.5G modem on compute die.
Interconnect : A20 Pro — planar metal routing, intra-die signaling. Kirin 9050 Pro — Cu-Cu bonding + high-density vertical TSVs; cross-die bandwidth 125 GB/s, critical-path latency reduced 30%+; 3D stacking shortens cache access paths.
Security : A20 Pro — Secure Enclave, dedicated security processor. Kirin 9050 Pro — dual-chip security: Mobile Security Processor (MSP) + independent security unit; quantum-resistant encryption.
Core Philosophy : A20 Pro — massive unified SLC-first on a large planar die. Kirin 9050 Pro — 3D logic-folding innovation — differentiated breakthrough for domestic chips.
Fundamental Architectural Divergence
A20 Pro represents the pinnacle of planar 2D optimization : leveraging 2nm to pack a 36 MB shared SLC into a ~98.7 mm² die, all modules communicating on one plane — simple design, controllable yield. Kirin 9050 Pro pursues a 3D stacking innovation path : splitting compute and cache into two vertically bonded dies, effectively placing cache "directly underneath" compute, drastically cutting cache access distance to compensate for process-node gap.
Cache Strategy Contrast
A20 Pro : monolithic unified SLC — CPU/GPU/ANE/ISP all share 36 MB, a long-standing Apple SoC strength. Kirin 9050 Pro : hierarchical private caches + 12 MB SLC; each CPU/NPU cluster owns dedicated L2/L3, all caches reside on the lower SRAM die, accessed via high-speed vertical interconnect.
Modem Integration
A20 Pro excludes cellular baseband (external discrete modem). Kirin 9050 Pro integrates Balong 5.5G modem directly on the compute die — a visible layout differentiator.
Generational Comparison: A18 Pro / A19 Pro / A20 Pro
Kurnal die shots show three generations sharing a highly consistent floorplan — progressive iteration with core differences in cache capacity, frequency, process, and transistor density.
Process : A18 Pro — TSMC N3E (3nm FinFET); A19 Pro — TSMC N3P (3nm enhanced FinFET); A20 Pro — TSMC N2 (2nm GAA nanosheet).
Die Area : A18 Pro — 105 mm²; A19 Pro — 98.69 mm²; A20 Pro — 98.72 mm².
CPU Spec : A18 Pro — 2P+4E, [email protected] GHz, [email protected] GHz; A19 Pro — 2P+4E, [email protected] GHz, [email protected] GHz; A20 Pro — 2P+4E, [email protected] GHz, [email protected] GHz.
GPU Spec : A18 Pro — 6-core, @1.47 GHz, HW ray-tracing; A19 Pro — 6-core, @1.62 GHz, HW ray-tracing; A20 Pro — 7-core, @1.62 GHz, HW ray-tracing.
SLC : A18 Pro — 24 MB (4×6 MB); A19 Pro — 32 MB (4×8 MB); A20 Pro — 36 MB (3×12 MB).
E-Core L2 : A18 Pro — 4 MiB; A19 Pro — 6 MiB; A20 Pro — 8 MiB.
P-Core L2 : all three generations — 32 MB (8+16+8).
Memory PHY : A18 Pro — 4×16-bit = 64-bit LPDDR5X; A19 Pro — 4×16-bit = 64-bit LPDDR5X; A20 Pro — 3×32-bit = 96-bit LPDDR5X.
ANE : A18 Pro — dual ANE, FP8 support; A19 Pro — dual ANE, FP8 enhanced; A20 Pro — dual 16-core ANE = 32 cores, FP8 throughput further improved.
Packaging : A18 Pro — traditional InFO-PoP (memory stacked atop SoC); A19 Pro — InFO-PoP; A20 Pro — WMCM — SoC and memory placed side-by-side.
Floorplan Skeleton : A18 Pro — GPU top, SLC center, CPU mid-low, ANE bottom-left, ISP/media bottom-right; A19 Pro — identical skeleton to A18, modules re-compacted; A20 Pro — skeleton retained; SLC blocks merged, GPU +1 core, memory PHY re-routed.
Evolution Summary
A18 → A19 : Same process family — "shrink area + grow cache + raise clocks." Floorplan topology nearly frozen; N3P density gains + layout re-optimization shrank die ~10% while expanding SLC/E-core caches and lifting frequencies. A high-density refinement iteration .
A19 → A20 : Jump to N2 GAA — "pile compute & bandwidth into a fixed area budget." Die area locked at ~98.7 mm²; new transistor budget spent on: +1 GPU core, larger SLC, bigger E-core L2, memory bus widened from 64-bit to 96-bit.
Across three generations Apple steadfastly pursues a large unified SLC strategy — SLC capacity climbing 24 → 32 → 36 MB, the most visible floorplan trait. All three retain 2P+4E 6-core CPU topology; big/little module positions on die barely move, only internal transistor-level microarchitecture evolves.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Architects' Tech Alliance
Sharing project experiences, insights into cutting-edge architectures, focusing on cloud computing, microservices, big data, hyper-convergence, storage, data protection, artificial intelligence, industry practices and solutions.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
