ByteDance’s 10‑Trillion‑Parameter Gamble: Inside China’s Largest AI Model Pre‑training

ByteDance is reportedly pre‑training a 10‑trillion‑parameter AI model that dwarfs domestic rivals, uses a Mixture‑of‑Experts architecture with low activation ratios, demands roughly 36 000 Blackwell GPUs and $2.5 billion in hardware, and deliberately avoids distilling competitor models, raising questions about China’s AI frontier timeline.

21CTO
21CTO
21CTO
ByteDance’s 10‑Trillion‑Parameter Gamble: Inside China’s Largest AI Model Pre‑training

According to the Financial Times, ByteDance is pre‑training a model with up to 10 trillion parameters, a scale that would surpass the current domestic leaders such as Moonshot AI’s Kimi K3 (2.8 trillion) and Meituan’s LongCat‑2.0/DeepSeek V4‑Pro (1.6 trillion) by roughly 3.5‑6×.

The article points out that raw parameter counts can be misleading because the model is expected to employ a Mixture‑of‑Experts (MoE) architecture. In an MoE system only a small fraction of parameters are active per forward pass: DeepSeek‑V3 activates about 5.5 % of its parameters, DeepSeek V4‑Pro reduces this to 3.1 %, and Kimi K3 routes merely 16 of its 896 expert networks.

Therefore, the headline‑grabbing “10 trillion” figure mainly reflects the memory bill rather than the compute bill; the actually active parameter count is estimated to lie between 200 billion and 500 billion.

Training such a model requires massive compute resources. With an anticipated token budget of 15‑40 trillion tokens, the total pre‑training FLOP cost is projected at 2 × 10²⁵‑1.2 × 10²⁶. ByteDance reportedly secured about 36 000 NVIDIA Blackwell GPUs from a Malaysian cloud provider, equivalent to roughly 500 GB200 racks, at an estimated hardware cost of $2.5 billion. The pre‑training phase is expected to last 3‑6 months, after which fine‑tuning and product release will follow.

The report also highlights ByteDance’s strategic decision to avoid “distilling” competitor models for over a year, opting instead for an independent research path. If the 10‑trillion‑parameter pre‑training succeeds, it would demonstrate ByteDance’s ability to independently develop frontier‑scale models without relying on external “teacher” models, challenging earlier forecasts that Chinese models would not reach this scale until 2027.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

Mixture of ExpertsByteDanceGPU ComputeAI scalingModel training costDistillation avoidance
21CTO
Written by

21CTO

21CTO (21CTO.com) offers developers community, training, and services, making it your go‑to learning and service platform.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.