Baidu Opens Tianchi Super Node Architecture Spec for High-Density AI Cabinets
Baidu releases the Tianchi Super Node system architecture design specification, covering cabinet, power, cooling, node, interconnect, and management modules with concrete specs for 32/64-GPU cabinets, hybrid air-liquid cooling, 1.5μs interconnect latency, and two-level BMC management to enable standardized, reusable high-density AI infrastructure.
Overview
Baidu has officially opened the Tianchi Super Node System Architecture Design Specification for download, providing a standardized, reusable reference architecture for high-density AI infrastructure. The specification documents design details and trade-offs validated at scale across six modules: cabinet, power supply, cooling, nodes, interconnect, and management.
Cabinet: Standard 46U Form Factor for Data Center Compatibility
The super node adopts a standard-width 46U cabinet that integrates compute, power, cooling, and interconnect modules within limited space while accommodating diverse data center deployment conditions. The specification defines cabinet dimensions, internal layout, pipeline and interface specifications for water, electricity, and network, plus blind-mate structural and interface designs to ensure simple, reliable installation and maintenance at high AI chip density.
Power Supply: 1% Efficiency Gain, Support for Hundreds of Kilowatts
With single-cabinet power reaching hundreds of kilowatts, traditional per-node power delivery cannot adapt to super-node density. The Tianchi design uses modular, configurable power architecture to cover the widest power range with minimal hardware variants, achieving a 1% power efficiency improvement and 15% increase in EDPP (Extreme Dynamic Pulse Power) support capability . The specification defines the complete power path—from facility-side input, cabinet-level conversion, to node-level consumption—and provides hardware layout, configuration, and redundancy architectures for 32-GPU and 64-GPU scales. It also addresses single-phase AC, DC, and three-phase AC facility power with corresponding hardware and cabling requirements.
Cooling: Hybrid Air-Liquid Architecture for Sustained Chip Performance
Cooling directly impacts both cabinet stability and sustained high-power AI chip output. The Tianchi super node employs a hybrid air-liquid cooling architecture : high-power components (AI chips, CPUs) use cold-plate liquid cooling to maintain peak performance, while low-power components use low-speed air cooling to reduce overall energy consumption. The specification details cooling designs for the cabinet, compute nodes, and switch nodes, with emphasis on Manifold, node cold plates, liquid coolant, and two-stage leak detection , balancing thermal efficiency, deployment cost, and operational safety.
Nodes: 1U Pluggable Design for Serviceability and Upgradability
Compute and network nodes are the smallest field-replaceable units. The Tianchi super node adopts a 1U per node pluggable design , unifying form factor and interface requirements for both compute and switch nodes. Nearby placement shortens the communication path between AI chips and CPUs. The specification defines front and rear window structures, power and cooling module requirements, and standardizes critical interfaces such as liquid-cooling quick-connects and high-speed connectors to reserve space for future AI chip and CPU adaptation, replacement, and upgrades.
Interconnect: Cable Tray Architecture for 1.5 μs Single-Hop Latency
Interconnect determines whether multi-GPU collaboration can be fully realized. In 32/64-GPU Scale-Up scenarios, the Tianchi super node achieves 1.5 μs single-hop latency via high-speed interconnect. A Cable Tray architecture provides stable physical lanes for high-speed signals. The specification constrains signal integrity from structure to components to assembly, detailing cable and connector selection, assembly methods, and impedance control requirements.
Management: Two-Level BMC with Unified Firmware Framework
Super nodes concentrate far more hardware than traditional servers—firmware count alone is dozens of times higher. The Tianchi super node implements a cabinet-level BMC plus node-level BMC hierarchical control plane, unified under a common firmware framework. This enables second-level metric collection and minute-level fault recovery . The specification defines the layered out-of-band management architecture from cabinet to node, and based on a unified hardware data model covers monitoring, fault isolation, and firmware lifecycle for power supplies, liquid cooling, motherboards, AI accelerator cards, NICs, and other components, giving the entire cabinet unified observability, controllability, and recoverability.
Open Reuse to Accelerate Industry Adoption
The specification encapsulates end-to-end practice from design and validation to volume deployment, incorporating collaboration experience from server vendors, component suppliers, and other ecosystem partners. By standardizing and publishing already-validated architectures, specifications, and interface requirements, Baidu aims to lower the entry barrier for super-node design and manufacturing, give data centers a reference for planning, deployment, and operations, and accelerate the普及 of high-density AI compute infrastructure.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Baidu Intelligent Cloud Tech Hub
We share the cloud tech topics you care about. Feel free to leave a message and tell us what you'd like to learn.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
