Why CPUs Are Back on Top: Inside HaiGuang’s Agent‑to‑Token Open Computing Architecture

The article explains HaiGuang’s Agent‑to‑Token open computing architecture, detailing how a dual‑chip CPU + DCU design enables end‑to‑end AI agent workflows, token generation, and secure, high‑performance compute through heterogeneous collaboration, open interconnects, and a full software stack.

Architects' Tech Alliance
Architects' Tech Alliance
Architects' Tech Alliance
Why CPUs Are Back on Top: Inside HaiGuang’s Agent‑to‑Token Open Computing Architecture

With the explosive growth of generative AI models and the scaling of Agent applications, the token (the smallest unit of AI information processing) has seen exponential usage, leading 2026 to be dubbed the “year of token economy.” HaiGuang introduced the “Agent to Token” open computing architecture at the 2026 China International Big Data Industry Expo, leveraging a CPU + DCU dual‑chip and open interconnect to create a full‑stack compute framework that closes the loop from business demand to Agent scheduling, token production, and result delivery.

What is the Agent‑to‑Token open computing architecture? In short, it is a full‑stack compute system built on HaiGuang’s CPU + DCU chips and open interconnect, spanning the entire workflow from Agent applications to token generation. The architecture focuses on AI workload needs, offering compute efficiency, security, and ecosystem enablement for high‑quality intelligent output across industries.

The architecture maps Agent tasks to CPU responsibilities (task orchestration, real‑time decisions, high‑frequency tool calls) and token production to DCU responsibilities (parallel batch decoding, KV‑cache acceleration, pre‑processing, and sequence support). This division allows the CPU to handle Agent‑centric workloads while the DCU accelerates token generation.

CPU’s role in the Agent‑to‑Token stack – The CPU acts as the system hub for Agent workloads, handling task scheduling, real‑time decision making, high‑frequency tool invocation, vector database and memory retrieval, sandbox isolation, and permission boundaries. HaiGuang achieves low latency (over 90% of Agent‑side processing) by moving from a traditional CPU : GPU 1:8 ratio toward a more balanced 1:4 or even 1:1 configuration, highlighting the resurgence of CPUs in AI.

Key CPU capabilities include high IPC/single‑core performance for data pre‑/post‑processing, many cores/multithreading for operator dispatch, and direct CPU‑DCU PCIe links with high channel count and bandwidth to support KV‑cache data transfer.

DCU’s role – The DCU handles parallel token generation, covering batch decoding, KV‑cache acceleration, pre‑processing, sequence support, and inference optimizations. It supports FP8/FP16/BF16 mixed‑precision compute, deep operator‑level tuning, and compiler optimizations to unlock hardware potential. For long‑context latency, DCU leverages DTK‑optimized FlashAttention and KV‑cache quantization to reduce pre‑compute overhead.

DCU also enables continuous batch processing, prefix‑cache reuse, and FP8 KV‑cache memory compression, dramatically raising single‑card concurrency limits. Its deep integration of sparse attention, KV‑cache quantization, and the proprietary HSL high‑speed interconnect allows linear scaling from a single card to tens of thousands of cards with low latency.

Four open capabilities

Compute openness : Heterogeneous, flexible CPU + DCU combinations, as well as CPU + xPU or CPU + DCU cloud‑edge configurations.

Interconnect openness : The HSL protocol enables low‑latency links across GPUs, CPUs, NICs, and SSDs, fostering a broad domestic interconnect ecosystem.

Security openness : Built‑in PSP (platform security processor) and CCP (cryptographic co‑processor) provide trusted boot, data protection, access control, remote attestation, and secure operations for the entire Agent‑to‑Token pipeline.

Software stack openness : DTK, DAS, and DAP expose a full suite of AI frameworks, compilers, runtimes, operator libraries, inference engines, and container orchestration, reducing model migration, development, and heterogeneous deployment costs.

DTK offers a mature compute library covering training and inference across AI4S scenarios, while DAS integrates over 2,000 operators supporting 100+ mainstream AI tools, already optimized for models such as GLM, DeepSeek, Kimi, and Qwen. DAP provides knowledge‑base and agent‑orchestration engines, with OpenDAS extensions and model repositories enabling OEMs and partners to integrate AI solutions easily.

Diagram
Diagram

In conclusion, the AI industry is shifting from a “training‑first” to an “inference‑first” paradigm, moving from models to agents. HaiGuang’s Agent‑to‑Token architecture systematically addresses this shift by reorganizing compute resources into a service‑oriented system that matches precise compute to each stage of Agent scheduling and token generation, redefining the metric of effective tokens over peak compute.

Architecture diagram
Architecture diagram
Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

AgentCPUTokenHeterogeneous computingAI computeOpen architectureDCUHSL
Architects' Tech Alliance
Written by

Architects' Tech Alliance

Sharing project experiences, insights into cutting-edge architectures, focusing on cloud computing, microservices, big data, hyper-convergence, storage, data protection, artificial intelligence, industry practices and solutions.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.