AMD Helios Rack-Scale AI System: 72-GPU Architecture, UALink Interconnect & 2026-2027 Roadmap
AMD's Helios rack integrates 72 MI455X GPUs (CDNA 5, 2nm) delivering 2.9 EFLOPS FP4 and 31 TB HBM4, using open UALink/UALoE interconnect with 102.4 Tb/s aggregate bandwidth, directly challenging NVIDIA NVL72 while outlining a roadmap to 256-GPU pods by 2027.
AMD announced the Helios rack-scale AI system in July 2026, positioning it as a direct competitor to NVIDIA's NVL72. A single Helios rack integrates 72 Instinct MI455X GPUs, 31 TB of HBM4 memory, and delivers 2.9 EFLOPS of FP4 inference throughput.
MI455X GPU Specifications
The MI455X is built on the CDNA 5 architecture using a 2 nm process with chiplet packaging, integrating 320 billion transistors. It is not sold standalone but exclusively as part of the Helios rack solution, which also includes EPYC 9006 series CPUs (codenamed "Venice") and Pensando networking. The rack provides 260 TB/s of expansion bandwidth and leverages the open UALink standard to unify the 72 GPUs into a single memory domain.
Comparison with NVIDIA Vera Rubin NVL72
FP4 compute: Helios 2.9 EFLOPS vs. Vera Rubin 3.6 EFLOPS (gap described as "not large").
Memory capacity: Helios 31 TB HBM4, 1.5× the NVL72's capacity.
Interconnect total bandwidth: Helios UALink matches Vera Rubin.
AMD Data Center Network Roadmap (2023–2027)
The roadmap splits into Scale-Up (intra-rack) and Scale-Out (inter-rack) networking:
2023–2025 (8-GPU era): Universal Baseboard (UBB) PCB interconnect, bandwidth rising from 7×128 GB/s to 7×153.6 GB/s, lane speed 32→38 Gb/s.
2026 (MI455X inflection): Logical GPU count jumps to 72 with a single-layer all-to-all topology (any GPU single-hop). Medium shifts to copper DAC + active retimers. Per-GPU bandwidth reaches 7×800 GB/s at 200 Gb/s lane speed. Dedicated switch ASICs provide 102.4 Tb/s aggregate bandwidth.
2027 (MI500X era): Scales to 256 GPUs in a 3-rack pod. Interconnect upgrades to co-packaged copper (CPC) or co-packaged optics (CPO). Per-GPU bandwidth targets 2,400 GB/s.
The article notes that "super-node" is AMD's core direction, moving from 8-GPU UBB to 72-GPU and 256-GPU super-node architectures to reduce cross-node communication overhead and improve large-model training efficiency, directly mirroring NVIDIA's NVL72 and NVL576 approaches.
Helios Rack Physical Layout
The rack contains 18 compute trays (9 top, 9 bottom) and 6 switch trays interleaved. Each compute tray holds 4 MI450 GPUs, forming a unified 72-GPU domain. Each switch tray houses two Tomahawk 6 102.4T Ethernet switches. The double-width mechanical design reserves space for future expansion to 144 GPUs without redesigning rack infrastructure.
Three-Layer Interconnect Architecture
CPU–GPU: Each MI455X connects via dedicated AMD Infinity Fabric to the CPU at 256 GB/s bidirectional, supporting cache-coherent access to GPU memory for KV-cache, data preprocessing, and multimodal workloads, avoiding PCIe copy overhead.
UALoE Scale-Up (intra-rack 72-GPU): Each MI455X exposes 36×400G UALoE links, cabled directly to the central switch trays through blind-mate copper. Handles shared memory, parameter sync, and MoE expert routing.
UALink Scale-Out (inter-rack cluster): Trays carry two custom NIC daughter cards integrating multiple AMD Pensando Vulcano 800 AI-NICs. Up to three 800G AI-NICs per GPU, each providing 600 GB/s bidirectional scale-out bandwidth. Rack-level aggregate reaches 43 TB/s over standard RDMA Ethernet for multi-rack cluster scaling.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Architects' Tech Alliance
Sharing project experiences, insights into cutting-edge architectures, focusing on cloud computing, microservices, big data, hyper-convergence, storage, data protection, artificial intelligence, industry practices and solutions.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
