Tagged articles

PD Separation

4 articles · Page 1 of 1
Machine Heart
Machine Heart
Jul 20, 2026 · Artificial Intelligence

Li Xiuhong on Cross-Cluster Heterogeneous PD Separation and Token Factory “Super Pipeline” at WAIC

The article details how 无问芯穹’s Agentic Infra strategy uses a cross‑cluster heterogeneous PD‑separation architecture (PDD) to cut first‑token latency by 51.5% and token cost by 37.5%, explains the bandwidth bottleneck of KV‑Cache transfer, introduces Decode‑side RadixCache and three‑stage handoff mechanisms, and shows a 37.5% BCR improvement that translates into roughly ten‑fold inference cost reduction.

Agentic InfraCross-ClusterHeterogeneous Computing
0 likes · 25 min read
Li Xiuhong on Cross-Cluster Heterogeneous PD Separation and Token Factory “Super Pipeline” at WAIC
DataFunSummit
DataFunSummit
Jun 21, 2026 · Artificial Intelligence

Unified Scheduling Optimization for xLLM in Complex Business Scenarios

This article analyzes how the xLLM open‑source LLM inference engine tackles the coexistence of multiple priority levels and strict SLO latency targets by introducing a dynamic, SLO‑aware batch scheduler and a PD‑separation architecture that improve throughput and SLO satisfaction across diverse workloads.

Hierarchical Block ManagerKV cachePD Separation
0 likes · 13 min read
Unified Scheduling Optimization for xLLM in Complex Business Scenarios
DataFunSummit
DataFunSummit
Mar 21, 2026 · Artificial Intelligence

How Slidebatching Revolutionizes LLM Inference Scheduling for Faster, More Efficient AI Services

The article examines the memory and latency challenges of 1750‑billion‑parameter LLM inference, introduces the xLLM framework’s Slidebatching and PD‑separation scheduling strategies, and details how these techniques achieve up to 35% system‑throughput gains and 52% SLO compliance improvements in real‑world multi‑priority workloads.

AI performanceLLMPD Separation
0 likes · 15 min read
How Slidebatching Revolutionizes LLM Inference Scheduling for Faster, More Efficient AI Services
Tencent Technical Engineering
Tencent Technical Engineering
Jan 23, 2026 · Artificial Intelligence

Unlocking AI Infra: Distributed Inference, PD Separation, TileLang, and Next‑Gen Agent Infrastructure

This article surveys the 2025 AI infrastructure landscape, covering distributed inference with PD‑separation, dynamic DOPD scheduling, AFD attention‑FFN disaggregation, high‑bandwidth cross‑machine communication libraries, the TileLang programming model, RL train‑inference decoupling via SeamlessFlow, and secure, low‑latency agent infra designs for future large‑scale models.

AI infrastructureAgent SystemsDistributed Inference
0 likes · 27 min read
Unlocking AI Infra: Distributed Inference, PD Separation, TileLang, and Next‑Gen Agent Infrastructure