Tagged articles

trillion-parameter model

3 articles · Page 1 of 1
Meituan Technology Team
Meituan Technology Team
Jul 1, 2026 · Artificial Intelligence

LongCat‑2.0: Training a Trillion‑Parameter Model on a Domestic 50k‑Card Cluster

Meituan’s LongCat‑2.0, a 1.6‑trillion‑parameter MoE model trained on a 50,000‑card domestic cluster, supports 1 M‑token context, uses Sparse Attention, zero‑compute experts and MOPD architecture, achieving over 1 T tokens/day throughput, 1.5× MFU efficiency, and top‑ranked scores on coding and agent benchmarks.

AI modelAgentic CodingLongCat-2.0
0 likes · 10 min read
LongCat‑2.0: Training a Trillion‑Parameter Model on a Domestic 50k‑Card Cluster
Data STUDIO
Data STUDIO
May 6, 2026 · Artificial Intelligence

DeepSeek V4 (Flash & Pro) Unveils Million‑Token Context and Trillion‑Parameter Inference

The April 24, 2026 release of DeepSeek V4 introduces Hybrid Attention (CSA/HCA), Manifold‑Constrained Hyper‑Connections, and the Muon optimizer, delivering 1 M‑token context windows, up to 1.6 T parameters, competitive benchmark scores against Claude and GPT, dramatically lower inference costs, and detailed deployment guidelines that expose both performance gains and practical challenges.

AI benchmarkingDeepSeek V4Hybrid Attention
0 likes · 17 min read
DeepSeek V4 (Flash & Pro) Unveils Million‑Token Context and Trillion‑Parameter Inference
Smart Sea Tide
Smart Sea Tide
Jan 28, 2026 · Artificial Intelligence

Alibaba Unveils Qwen3‑Max‑Thinking: Trillion‑Parameter Model Joins Global AI Elite

Alibaba's Tongyi team released the Qwen3‑Max‑Thinking model, surpassing one trillion parameters and 36 T tokens of pre‑training, introducing adaptive tool‑calling and test‑time expansion techniques that boost benchmark scores across 19 tests, positioning it alongside top international models and now available via app, web, and API.

AI model comparisonQwen3-Max-Thinkingadaptive tool calling
0 likes · 5 min read
Alibaba Unveils Qwen3‑Max‑Thinking: Trillion‑Parameter Model Joins Global AI Elite