Huawei Atlas 960E SuperPoD: First NPO Super Node for 10-Trillion-Parameter Models
Huawei unveils Atlas 960E SuperPoD, the industry's first super node using Near Package Optics (NPO) technology, scaling to 4,096 Ascend 960 chips delivering 8 EFLOPS FP8 and 16 EFLOPS FP4 for trillion-parameter model training, with 5,500 Hi-ONE optical engines replacing 48,000 modules to cut 550 kW power and achieve 99.8% availability.
At the 2026 Huawei Connect conference on September 17, Huawei officially launched the Ascend 960 Super Node , marketed as the Atlas 960E SuperPoD . This system is the industry's first super node to adopt Near Package Optics (NPO) technology , built on the "Lingqu UnifiedBus + Hi-ONE" architecture to form a scalable super-node interconnect system designed to accelerate training and inference of models with ten-trillion parameters .
1. Atlas 960E Super Node Core Specifications and Performance
Compute and Scale: A single super node scales to 4,096 cards , delivering 8 EFLOPS FP8 and 16 EFLOPS FP4 compute throughput.
Architecture Design: Based on "Lingqu + Hi-ONE" , employing an orthogonal architecture with full liquid cooling and supporting unified memory addressing .
Optical Interconnect Breakthrough: Uses 5,500 Hi-ONE optical engines to replace the 48,000 800G optical modules otherwise required. This reduces power consumption by over 550 kW and doubles mean time between failures, achieving 99.8% availability .
Cluster Capability: Multiple super nodes interconnect via Lingqu network or RoCE to build clusters up to 512,000 cards , with a roadmap to 1 million cards in Ascend super-node clusters.
2. Ascend 950 SuperPoD Progress
Alongside the 960 launch, Huawei disclosed that the previous-generation Atlas 950 SuperPoD has entered large-scale commercial deployment . Its physical form comprises compute cabinets and Lingqu interconnect cabinets . A single cabinet supports 64 super nodes; the system scales to 1,024 super nodes. Over 1,000 Atlas 950 SuperPoD units have been deployed commercially , and the Ascend super node is noted as the only domestic super node that has trained SOTA (state-of-the-art) models .
3. Ascend Processor Release Schedule and Microarchitecture Details
Ascend 960DT — Ready Q1 2027 (three quarters ahead of plan)
Microarchitecture: SIMD/SIMT
Compute: 2 PFLOPS (FP8) / 4 PFLOPS (FP4)
Data formats: FP32, HF32, FP16, HiF4, etc.
Interconnect bandwidth: 2.2 TB/s
Memory: Capacity doubled to 288 GB , bandwidth 9.6 TB/s
Ascend 960PR — Ready Q3 2027 (one quarter ahead of plan)
Key change: FP4 compute raised to 8 PFLOPS
Trade-off: Memory capacity reduced to 192 GB , bandwidth to 2.4 TB/s
Ascend 970 — Planned 2028
Microarchitecture: SIMD/SIMT
Compute: 3.6 PFLOPS (FP8) / 14 PFLOPS (FP4)
Data formats: FP32, HF32, FP16, BF16, FP8, MXFP8, HiF8, HiF4
Memory: 288 GB capacity, interconnect bandwidth 4.4 TB/s
Ascend 980 — Planned 2029
Microarchitecture: SIMD/SIMT
Compute: 7.2 PFLOPS (FP8) / 28 PFLOPS (FP4)
Data formats: FP32, HF32, FP16, BF16, FP8, HiF4
Memory: 384 GB capacity, interconnect bandwidth 8 TB/s
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Architects' Tech Alliance
Sharing project experiences, insights into cutting-edge architectures, focusing on cloud computing, microservices, big data, hyper-convergence, storage, data protection, artificial intelligence, industry practices and solutions.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
