Unveiling Scaling Laws for Mixture-of-Experts Diffusion Language Models: LLaDA MoE v2
The paper systematically derives scaling laws for mixture-of-experts diffusion language models, trains a 30B-parameter LLaDA MoE v2 model guided by these laws, and demonstrates that it matches or exceeds state-of-the-art autoregressive models while using far less compute.
