ByteDance’s 10‑Trillion‑Parameter Gamble: Inside China’s Largest AI Model Pre‑training
ByteDance is reportedly pre‑training a 10‑trillion‑parameter AI model that dwarfs domestic rivals, uses a Mixture‑of‑Experts architecture with low activation ratios, demands roughly 36 000 Blackwell GPUs and $2.5 billion in hardware, and deliberately avoids distilling competitor models, raising questions about China’s AI frontier timeline.
According to the Financial Times, ByteDance is pre‑training a model with up to 10 trillion parameters, a scale that would surpass the current domestic leaders such as Moonshot AI’s Kimi K3 (2.8 trillion) and Meituan’s LongCat‑2.0/DeepSeek V4‑Pro (1.6 trillion) by roughly 3.5‑6×.
The article points out that raw parameter counts can be misleading because the model is expected to employ a Mixture‑of‑Experts (MoE) architecture. In an MoE system only a small fraction of parameters are active per forward pass: DeepSeek‑V3 activates about 5.5 % of its parameters, DeepSeek V4‑Pro reduces this to 3.1 %, and Kimi K3 routes merely 16 of its 896 expert networks.
Therefore, the headline‑grabbing “10 trillion” figure mainly reflects the memory bill rather than the compute bill; the actually active parameter count is estimated to lie between 200 billion and 500 billion.
Training such a model requires massive compute resources. With an anticipated token budget of 15‑40 trillion tokens, the total pre‑training FLOP cost is projected at 2 × 10²⁵‑1.2 × 10²⁶. ByteDance reportedly secured about 36 000 NVIDIA Blackwell GPUs from a Malaysian cloud provider, equivalent to roughly 500 GB200 racks, at an estimated hardware cost of $2.5 billion. The pre‑training phase is expected to last 3‑6 months, after which fine‑tuning and product release will follow.
The report also highlights ByteDance’s strategic decision to avoid “distilling” competitor models for over a year, opting instead for an independent research path. If the 10‑trillion‑parameter pre‑training succeeds, it would demonstrate ByteDance’s ability to independently develop frontier‑scale models without relying on external “teacher” models, challenging earlier forecasts that Chinese models would not reach this scale until 2027.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
21CTO
21CTO (21CTO.com) offers developers community, training, and services, making it your go‑to learning and service platform.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
