Alibaba Panjiu AL64 Super Node: 500K-GPU Scale Architecture with Zhenwu V900 Chip
Alibaba Cloud's Panjiu AL64 super node integrates the Zhenwu V900 AI chip, ICN Switch, Panmai smart NIC, and Zhenyue SSD controller to enable 500,000-GPU cluster scaling via a three-tier interconnect architecture and liquid-cooled orthogonal design.
Panjiu Super Node Product Roadmap (2026–2028)
The 2026 Yunqi Conference revealed a four-stage roadmap for Panjiu super nodes:
2026 – AL128 (Gen 1): 128 AI accelerator cards per cabinet, 350 kW power density requiring liquid cooling, 150 ns ultra-low latency, petabyte-scale Scale-Up bandwidth, supports Zhenwu M890.
2027 – AL128 (Upgrade): Retains 128-card form factor, adds 800 V HVDC (high-voltage DC power) for improved supply efficiency, supports Zhenwu V900 and other GPUs.
2027 – AL64 (Distributed Architecture): 64 cards per cabinet, distributed cabinet deployment (unlike prior high-density single cabinets), Scale-Up to 1,024-card clusters; adopts SNPO all-optical interconnect to reduce long-distance communication latency, supports Zhenwu V900.
2028 – AL144 (Ultimate Form): 144 cards per cabinet, 650 kW (double power), Scale-Up to 10,368 cards (massive cluster expansion), doubled Scale-Up bandwidth, vertical micro-channel + liquid metal cooling (advanced thermal), native 800 V HVDC, supports Zhenwu J900 and other GPUs.
Panjiu AI-Native New-Generation Super Node Architecture
The Panjiu AL64 Zhenwu V900 super node, paired with Panjiu AI Infra 3.0, targets Agent-era large-model training and ultra-large-cluster inference. It achieves full-stack integration of self-developed compute, interconnect, NIC, and storage controller chips, natively supports wide-area super nodes, and scales a single AI cluster to 500,000 cards. Key architectural features:
Orthogonal Decoupled Architecture: CPU head and XPU tail separated; three blind-mate modular design enables compute, switch, and storage module replacement without cabinet removal, cutting maintenance time from hours to minutes.
Single-Node Capacity: 64 Zhenwu V900 cards; Scale-Up non-convergent expansion to 1,024 cards within a single super-node cabinet domain.
Three-Layer Interconnect Fabric
Scale-Up Domain (Intra-Cabinet): ICN Switch 2.0, thousand-card full bandwidth, native memory semantics, unified addressing.
Scale-Out Domain (Inter-Cabinet): Panmai 920 smart NIC + high-speed optical network for horizontal super-node expansion.
Wide-Area DCN Domain (Multi-DC Cross-Region): Alibaba Cloud next-gen intelligent computing network enabling wide-area super nodes, cluster ceiling of 500,000 cards.
Thermal & Power
Full liquid cooling; single compute module up to 3,500 W; cabinet supports high-power high-density deployment.
Hardware Interfaces
SNPO near-package optical interconnect, 51.2 Tbps optical switching, 64 × 800G optical ports; copper/optical shared socket, compatible with future 224G/448G evolution.
Deep Integration of Four Self-Developed Core Chips
The new Panjiu super node fuses T-Head's full-stack custom silicon:
Zhenwu V900 AI Accelerator: Next-gen train/inference unified chip; 3× comprehensive performance vs. previous M890 flagship; 216 GB VRAM; 1,200 GB/s inter-chip bandwidth; native FP8/FP4 low-precision compute; targets trillion-to-10-trillion parameter model training and inference; slated for Q1 2027 volume production.
ICN Switch Interconnect Chip: Custom high-speed interconnect switch; handles thousand-card full-bandwidth intra-super-node communication; reduces cross-card latency.
Panmai Smart NIC: Offloads Scale-Out network protocol processing from CPU to NIC, drastically lowering host overhead and boosting multi-node communication efficiency for large-scale distributed training.
Zhenyue SSD Controller: Custom storage controller optimized for AI workloads — KV-Cache, model weights, dataset high-speed read/write; bridges compute and storage paths, alleviating large-model inference storage bottlenecks.
Compared with the previous Panjiu AL128 (M890) which already validated 2-trillion-parameter model training, the new Panjiu + Zhenwu V900 pushes the model-scale ceiling higher while improving cluster compute utilization.
T-Head Zhenwu V900 AI Chip Details
The conference underscored a shift from "single-point compute" to "full-stack system" competition. T-Head committed to a "one generation per year" cadence: 2024 Zhenwu 810E (96 GB VRAM), 2026 M890 (144 GB VRAM), 2027 Q1 Zhenwu V900 (216 GB VRAM). Zhenwu V900 is built on a custom parallel compute architecture, delivers 216 GB VRAM, 1,200 GB/s inter-chip bandwidth, native FP8/FP4 support, raising compute density while significantly cutting inference cost.
Wide-area super nodes are not the endpoint. Post-V900, Zhenwu J900 will further break compute ceilings, and Alibaba Cloud's intelligent computing network will continue evolving to enhance cross-region wide-area compute collaboration.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Architects' Tech Alliance
Sharing project experiences, insights into cutting-edge architectures, focusing on cloud computing, microservices, big data, hyper-convergence, storage, data protection, artificial intelligence, industry practices and solutions.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
