Biren's 1024-GPU Super Node: NPO Optical Interconnect & Chiplet Architecture Explained

Biren Technology details a three-tier AI super node architecture scaling to 1024 GPUs via NPO optical interconnect and the BLink 2.0 protocol, built on its Chiplet-based BR20X GPUs with native FP8/FP4 support for trillion-parameter model training.

Architects' Tech Alliance
Architects' Tech Alliance
Architects' Tech Alliance
Biren's 1024-GPU Super Node: NPO Optical Interconnect & Chiplet Architecture Explained

BR20X Series GPU: Next-Generation Training Chip

Biren's next-generation flagship training chip, the BR20X series GPU, introduces native support for FP8 and FP4 low-precision formats to boost large-model training and inference efficiency. Compared with the first generation, it features larger-capacity, higher-speed memory and higher interconnect bandwidth to alleviate data-movement bottlenecks.

Interconnect & Scale-Up

The BR20X series integrates super-node interconnect capability natively, relying on the proprietary BLink 2.0 super-node interconnect protocol to connect up to 1024 GPUs in a single scale-up domain. The accompanying system design, "LightSphere X" , adopts NPO (Near-Package Optics) optical interconnect and a distributed decoupled architecture .

Memory-semantic interconnect: Enables up to 1024 GPUs to share a single memory space, operating as one "super GPU."

In-network computing: Offloads mathematical operations from communication to the switch, reducing GPU compute load.

Intelligent congestion control & link self-healing: Multi-layer protection from physical to framework layer to prevent network congestion and ensure stability for large-scale training jobs.

Three-Tier Super Node Product Matrix

Based on the BR2xx chips and BLink 2.0, Biren defines three product tiers:

Standard Server Super Node — Max scale: 16 GPUs; Interconnect: Electrical; Target scenario: SMEs, hundred-billion-parameter models.

High-Density Cabinet Super Node — Max scale: 128 GPUs; Interconnect: Electrical; Target scenario: Scaled deployment of hundred-billion to trillion-parameter models.

Distributed Decoupled Architecture Super Node — Max scale: 1024 GPUs; Interconnect: NPO Optical; Target scenario: Large customers, ten-trillion-parameter models; NPO optical interconnect and large-scale expansion suited for hyperscale scenarios.

1024-Card Distributed Decoupled Architecture (NPO Optical)

The flagship 1024-card solution physically separates GPU nodes from switch nodes, allowing "Lego-like" expansion to 1024 cards. It is based on the next-generation BR20X series GPU , with formal launch expected in H2 2026 , while large-scale commercial deployment of NPO optical interconnect is projected for 2028 . The roadmap also mentions BR30X series planned through 2028 .

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

large model trainingChiplet architectureBiren TechnologyAI super nodeBLink 2.0BR20X GPUFP8/FP4NPO optical interconnect
Architects' Tech Alliance
Written by

Architects' Tech Alliance

Sharing project experiences, insights into cutting-edge architectures, focusing on cloud computing, microservices, big data, hyper-convergence, storage, data protection, artificial intelligence, industry practices and solutions.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.