Baidu Intelligent Cloud Tech Hub
Author

Baidu Intelligent Cloud Tech Hub

We share the cloud tech topics you care about. Feel free to leave a message and tell us what you'd like to learn.

147
Articles
0
Likes
799
Views
0
Comments
Recent Articles

Latest from Baidu Intelligent Cloud Tech Hub

100 recent articles max
Baidu Intelligent Cloud Tech Hub
Baidu Intelligent Cloud Tech Hub
Sep 3, 2026 · Cloud Computing

Baidu Opens Tianchi Supernode Reference Architecture to Accelerate Industry Adoption

Baidu Intelligent Cloud has opened its Tianchi supernode reference architecture, a production-validated design for high-density AI compute clusters, to lower barriers for industry-wide adoption by providing a reusable blueprint covering interconnect, power, cooling, and maintenance principles proven across hundreds of cabinets running trillion-parameter models.

AI InfrastructureBaidudata center
0 likes · 17 min read
Baidu Opens Tianchi Supernode Reference Architecture to Accelerate Industry Adoption
Baidu Intelligent Cloud Tech Hub
Baidu Intelligent Cloud Tech Hub
Aug 27, 2026 · Big Data

PFS L3 Architecture Deep Dive: Redesigning Parallel File Storage for AI Production Workloads

Baidu's PFS L3 rearchitects parallel file storage for AI production loads with a hybrid data engine (Aries + BlockServer), scalable metadata foundation (MetaDB), high-performance client (RapidFC), RDMA network tuning, transparent tiering to object storage, and AgenticFS for multi-tenant Agent workloads, proven in RL training saving 30k GPU-hours weekly.

AI storageAgenticFSAries
0 likes · 29 min read
PFS L3 Architecture Deep Dive: Redesigning Parallel File Storage for AI Production Workloads
Baidu Intelligent Cloud Tech Hub
Baidu Intelligent Cloud Tech Hub
Aug 21, 2026 · Artificial Intelligence

Achieving 2.37× Faster OpenPI PyTorch Training with RLinf and Baidu Baige

By first aligning Pi0.5’s precision between JAX and PyTorch and then applying a full‑stack AI Infra overhaul—including data‑fetch slot pruning, uint8 H2D transfers, torch.compile fusion, per‑block FSDP restructuring and communication prefetch—the team raised OpenPI PyTorch training throughput from 74.2 sps to 175.6 sps (2.37×) while preserving accuracy and achieving over 91% scaling efficiency on 256 GPUs.

Baidu BaigeDistributed TrainingOpenPI
0 likes · 19 min read
Achieving 2.37× Faster OpenPI PyTorch Training with RLinf and Baidu Baige
Baidu Intelligent Cloud Tech Hub
Baidu Intelligent Cloud Tech Hub
Aug 17, 2026 · Information Security

Building a Secure Agent Framework: Lessons from OpenAI and Anthropic Risks

Recent OpenAI and Anthropic incidents reveal how unchecked AI agents can escape sandbox limits, prompting a detailed analysis that shows agents’ risks evolve step‑by‑step and proposes a security framework—defining what agents want, what they can do, and establishing comprehensive governance across the task lifecycle.

AI AgentsAgent GovernanceSecurity
0 likes · 12 min read
Building a Secure Agent Framework: Lessons from OpenAI and Anthropic Risks
Baidu Intelligent Cloud Tech Hub
Baidu Intelligent Cloud Tech Hub
Aug 7, 2026 · Cloud Computing

How Baidu Cloud’s Sonata Achieves Predictable Performance Across Multi‑Vendor, Multi‑Generation NICs (NSDI’27)

The paper presents Sonata, a software‑centric vSwitch data plane that decouples from specific NICs, enabling a single implementation to handle heterogeneous, multi‑generation hardware while delivering 2.13‑2.51× software CPS gains and an additional 4.3‑11.5× boost with hardware offload, all validated on hundreds of thousands of production servers.

Cloud NetworkingRobustnessSoftware-Centric
0 likes · 12 min read
How Baidu Cloud’s Sonata Achieves Predictable Performance Across Multi‑Vendor, Multi‑Generation NICs (NSDI’27)
Baidu Intelligent Cloud Tech Hub
Baidu Intelligent Cloud Tech Hub
Jul 31, 2026 · Databases

MetaDB Enables Databases to Truly Understand File Systems – Accepted to NSDI’27

MetaDB, the unified metadata layer for Baidu Cloud Storage, was accepted to NSDI'27 after more than five years of production use, delivering up to 10.7× higher same‑directory throughput, halving the required metadata machines, and addressing three fundamental shortcomings of generic distributed databases in handling file‑system metadata.

MetaDBNSDI 2027cloud storage
0 likes · 10 min read
MetaDB Enables Databases to Truly Understand File Systems – Accepted to NSDI’27
Baidu Intelligent Cloud Tech Hub
Baidu Intelligent Cloud Tech Hub
Jul 30, 2026 · Artificial Intelligence

Boosting Cosmos‑3 Training Throughput by 35% on the Same Budget with Heterogeneous GPU Clusters

By decoupling VAE encoding from the main training GPU, redesigning the data flow into a fully asynchronous pipeline, and applying dynamic scheduling, multi‑level caching, and custom operators, the Cosmos3‑Nano‑Policy‑DROID system achieves a 2.1× overall throughput with only 1.5× the hardware cost, delivering roughly 35% more training throughput for the same budget.

AI InfraGPU ClusterTraining Optimization
0 likes · 18 min read
Boosting Cosmos‑3 Training Throughput by 35% on the Same Budget with Heterogeneous GPU Clusters
Baidu Intelligent Cloud Tech Hub
Baidu Intelligent Cloud Tech Hub
Jul 20, 2026 · Cloud Native

Redesigning Parallel File Storage for AI Production: Introducing Baidu PFS L3

As AI moves from isolated model training to full‑scale production with mixed workloads, Baidu's new PFS L3 offers elastic, cloud‑native parallel file storage that delivers tens of GB/s throughput, millions of IOPS, sub‑millisecond latency, and automated data lifecycle management to meet the evolving demands of modern AI platforms.

AI InfrastructureBaidu CloudPFS L3
0 likes · 12 min read
Redesigning Parallel File Storage for AI Production: Introducing Baidu PFS L3
Baidu Intelligent Cloud Tech Hub
Baidu Intelligent Cloud Tech Hub
Jun 29, 2026 · Artificial Intelligence

How to Maximize Cosmos‑3 Training Throughput Without NVLink Using AI Infra Optimizations

By applying systematic AI Infra engineering—optimizing data loading, I/O pipelines, activation checkpointing, torch.compile, and multi‑node scaling—the Cosmos‑3‑Nano‑Policy‑DROID model achieved an 89× faster startup, 99.3% higher single‑node throughput, 0.42 MFU, and 98.3% scaling efficiency across 12 nodes, all without NVLink.

AI InfraActivation CheckpointingCosmos 3
0 likes · 13 min read
How to Maximize Cosmos‑3 Training Throughput Without NVLink Using AI Infra Optimizations
Baidu Intelligent Cloud Tech Hub
Baidu Intelligent Cloud Tech Hub
Jun 17, 2026 · Artificial Intelligence

The Multimodal Model Battlefield Is Going Rogue – LoongForge’s ‘Dark Arts’ Framework

Facing mounting challenges of heterogeneous models, data, and hardware in multimodal training, Baidu’s open‑source LoongForge framework unifies LLM, VLM, VLA and diffusion workloads, delivering 1.15‑2.31× speedups and over 5× gains for DSA models while scaling linearly across thousands of GPUs and Kunlun XPU cards.

GPUKunlun XPULoongForge
0 likes · 8 min read
The Multimodal Model Battlefield Is Going Rogue – LoongForge’s ‘Dark Arts’ Framework