Cloud Computing 10 min read

Baidu Opens Tianchi Super Node Architecture Spec for High-Density AI Cabinets

Baidu releases the Tianchi Super Node system architecture design specification, covering cabinet, power, cooling, node, interconnect, and management modules with concrete specs for 32/64-GPU cabinets, hybrid air-liquid cooling, 1.5μs interconnect latency, and two-level BMC management to enable standardized, reusable high-density AI infrastructure.

Baidu Intelligent Cloud Tech Hub
Baidu Intelligent Cloud Tech Hub
Baidu Intelligent Cloud Tech Hub
Baidu Opens Tianchi Super Node Architecture Spec for High-Density AI Cabinets

Overview

Baidu has officially opened the Tianchi Super Node System Architecture Design Specification for download, providing a standardized, reusable reference architecture for high-density AI infrastructure. The specification documents design details and trade-offs validated at scale across six modules: cabinet, power supply, cooling, nodes, interconnect, and management.

Cabinet: Standard 46U Form Factor for Data Center Compatibility

The super node adopts a standard-width 46U cabinet that integrates compute, power, cooling, and interconnect modules within limited space while accommodating diverse data center deployment conditions. The specification defines cabinet dimensions, internal layout, pipeline and interface specifications for water, electricity, and network, plus blind-mate structural and interface designs to ensure simple, reliable installation and maintenance at high AI chip density.

Cabinet layout diagram
Cabinet layout diagram

Power Supply: 1% Efficiency Gain, Support for Hundreds of Kilowatts

With single-cabinet power reaching hundreds of kilowatts, traditional per-node power delivery cannot adapt to super-node density. The Tianchi design uses modular, configurable power architecture to cover the widest power range with minimal hardware variants, achieving a 1% power efficiency improvement and 15% increase in EDPP (Extreme Dynamic Pulse Power) support capability . The specification defines the complete power path—from facility-side input, cabinet-level conversion, to node-level consumption—and provides hardware layout, configuration, and redundancy architectures for 32-GPU and 64-GPU scales. It also addresses single-phase AC, DC, and three-phase AC facility power with corresponding hardware and cabling requirements.

Power architecture diagram
Power architecture diagram

Cooling: Hybrid Air-Liquid Architecture for Sustained Chip Performance

Cooling directly impacts both cabinet stability and sustained high-power AI chip output. The Tianchi super node employs a hybrid air-liquid cooling architecture : high-power components (AI chips, CPUs) use cold-plate liquid cooling to maintain peak performance, while low-power components use low-speed air cooling to reduce overall energy consumption. The specification details cooling designs for the cabinet, compute nodes, and switch nodes, with emphasis on Manifold, node cold plates, liquid coolant, and two-stage leak detection , balancing thermal efficiency, deployment cost, and operational safety.

Cooling architecture diagram
Cooling architecture diagram

Nodes: 1U Pluggable Design for Serviceability and Upgradability

Compute and network nodes are the smallest field-replaceable units. The Tianchi super node adopts a 1U per node pluggable design , unifying form factor and interface requirements for both compute and switch nodes. Nearby placement shortens the communication path between AI chips and CPUs. The specification defines front and rear window structures, power and cooling module requirements, and standardizes critical interfaces such as liquid-cooling quick-connects and high-speed connectors to reserve space for future AI chip and CPU adaptation, replacement, and upgrades.

Interconnect: Cable Tray Architecture for 1.5 μs Single-Hop Latency

Interconnect determines whether multi-GPU collaboration can be fully realized. In 32/64-GPU Scale-Up scenarios, the Tianchi super node achieves 1.5 μs single-hop latency via high-speed interconnect. A Cable Tray architecture provides stable physical lanes for high-speed signals. The specification constrains signal integrity from structure to components to assembly, detailing cable and connector selection, assembly methods, and impedance control requirements.

Management: Two-Level BMC with Unified Firmware Framework

Super nodes concentrate far more hardware than traditional servers—firmware count alone is dozens of times higher. The Tianchi super node implements a cabinet-level BMC plus node-level BMC hierarchical control plane, unified under a common firmware framework. This enables second-level metric collection and minute-level fault recovery . The specification defines the layered out-of-band management architecture from cabinet to node, and based on a unified hardware data model covers monitoring, fault isolation, and firmware lifecycle for power supplies, liquid cooling, motherboards, AI accelerator cards, NICs, and other components, giving the entire cabinet unified observability, controllability, and recoverability.

Open Reuse to Accelerate Industry Adoption

The specification encapsulates end-to-end practice from design and validation to volume deployment, incorporating collaboration experience from server vendors, component suppliers, and other ecosystem partners. By standardizing and publishing already-validated architectures, specifications, and interface requirements, Baidu aims to lower the entry barrier for super-node design and manufacturing, give data centers a reference for planning, deployment, and operations, and accelerate the普及 of high-density AI compute infrastructure.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

System ArchitectureAI InfrastructureLiquid CoolingPower DistributionInterconnectHigh-Density ComputingBMC ManagementTianchi Super Node
Baidu Intelligent Cloud Tech Hub
Written by

Baidu Intelligent Cloud Tech Hub

We share the cloud tech topics you care about. Feel free to leave a message and tell us what you'd like to learn.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.