ByteDance’s 10‑Trillion‑Parameter Gamble: Inside China’s Largest AI Model Pre‑training
ByteDance is reportedly pre‑training a 10‑trillion‑parameter AI model that dwarfs domestic rivals, uses a Mixture‑of‑Experts architecture with low activation ratios, demands roughly 36 000 Blackwell GPUs and $2.5 billion in hardware, and deliberately avoids distilling competitor models, raising questions about China’s AI frontier timeline.
