Tagged articles

scaling efficiency

2 articles · Page 1 of 1
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jul 28, 2026 · Artificial Intelligence

The Most Valuable Innovations in Kimi K3’s 47‑Page Technical Report

Kimi K3 scales to 2.8 T parameters, 104 B activation parameters and a 1 M‑token context by introducing KDA‑based compressed sequence state, AttnRes for depth‑wise residual selection, Stable LatentMoE for efficient expert routing, partial‑rollout reinforcement learning and Firecracker micro‑VM sandboxing, achieving roughly 2.5× scaling efficiency and notable benchmark gains over competing models.

Agent infrastructureAttnResKDA
0 likes · 17 min read
The Most Valuable Innovations in Kimi K3’s 47‑Page Technical Report
Baidu Intelligent Cloud Tech Hub
Baidu Intelligent Cloud Tech Hub
Jun 29, 2026 · Artificial Intelligence

How to Maximize Cosmos‑3 Training Throughput Without NVLink Using AI Infra Optimizations

By applying systematic AI Infra engineering—optimizing data loading, I/O pipelines, activation checkpointing, torch.compile, and multi‑node scaling—the Cosmos‑3‑Nano‑Policy‑DROID model achieved an 89× faster startup, 99.3% higher single‑node throughput, 0.42 MFU, and 98.3% scaling efficiency across 12 nodes, all without NVLink.

AI InfraActivation CheckpointingCosmos 3
0 likes · 13 min read
How to Maximize Cosmos‑3 Training Throughput Without NVLink Using AI Infra Optimizations