LongCat‑2.0: Training a Trillion‑Parameter Model on a Domestic 50k‑Card Cluster
Meituan’s LongCat‑2.0, a 1.6‑trillion‑parameter MoE model trained on a 50,000‑card domestic cluster, supports 1 M‑token context, uses Sparse Attention, zero‑compute experts and MOPD architecture, achieving over 1 T tokens/day throughput, 1.5× MFU efficiency, and top‑ranked scores on coding and agent benchmarks.
