Biren's 1024-GPU Super Node: NPO Optical Interconnect & Chiplet Architecture Explained
Biren Technology details a three-tier AI super node architecture scaling to 1024 GPUs via NPO optical interconnect and the BLink 2.0 protocol, built on its Chiplet-based BR20X GPUs with native FP8/FP4 support for trillion-parameter model training.
BR20X Series GPU: Next-Generation Training Chip
Biren's next-generation flagship training chip, the BR20X series GPU, introduces native support for FP8 and FP4 low-precision formats to boost large-model training and inference efficiency. Compared with the first generation, it features larger-capacity, higher-speed memory and higher interconnect bandwidth to alleviate data-movement bottlenecks.
Interconnect & Scale-Up
The BR20X series integrates super-node interconnect capability natively, relying on the proprietary BLink 2.0 super-node interconnect protocol to connect up to 1024 GPUs in a single scale-up domain. The accompanying system design, "LightSphere X" , adopts NPO (Near-Package Optics) optical interconnect and a distributed decoupled architecture .
Memory-semantic interconnect: Enables up to 1024 GPUs to share a single memory space, operating as one "super GPU."
In-network computing: Offloads mathematical operations from communication to the switch, reducing GPU compute load.
Intelligent congestion control & link self-healing: Multi-layer protection from physical to framework layer to prevent network congestion and ensure stability for large-scale training jobs.
Three-Tier Super Node Product Matrix
Based on the BR2xx chips and BLink 2.0, Biren defines three product tiers:
Standard Server Super Node — Max scale: 16 GPUs; Interconnect: Electrical; Target scenario: SMEs, hundred-billion-parameter models.
High-Density Cabinet Super Node — Max scale: 128 GPUs; Interconnect: Electrical; Target scenario: Scaled deployment of hundred-billion to trillion-parameter models.
Distributed Decoupled Architecture Super Node — Max scale: 1024 GPUs; Interconnect: NPO Optical; Target scenario: Large customers, ten-trillion-parameter models; NPO optical interconnect and large-scale expansion suited for hyperscale scenarios.
1024-Card Distributed Decoupled Architecture (NPO Optical)
The flagship 1024-card solution physically separates GPU nodes from switch nodes, allowing "Lego-like" expansion to 1024 cards. It is based on the next-generation BR20X series GPU , with formal launch expected in H2 2026 , while large-scale commercial deployment of NPO optical interconnect is projected for 2028 . The roadmap also mentions BR30X series planned through 2028 .
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Architects' Tech Alliance
Sharing project experiences, insights into cutting-edge architectures, focusing on cloud computing, microservices, big data, hyper-convergence, storage, data protection, artificial intelligence, industry practices and solutions.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
