Tagged articles

Attention Residuals

4 articles · Page 1 of 1
AI Engineering
AI Engineering
Aug 2, 2026 · Artificial Intelligence

Kimi K3 Technical Report Reveals Answers to Key Architecture Questions

The 47‑page Kimi K3 technical report, released on July 27, details the 2.8‑trillion‑parameter model’s novel LatentMoE, SiTU‑GLU, Quantile Balancing, Attention Residuals, and full‑stack NoPE design, explains how these solve activation‑explosion and load‑imbalance problems, and provides open‑source code for inference and agentic RL.

Attention ResidualsKimi K3LatentMoE
0 likes · 12 min read
Kimi K3 Technical Report Reveals Answers to Key Architecture Questions
Machine Heart
Machine Heart
Apr 3, 2026 · Artificial Intelligence

Kimi’s ‘Option Time Machine’: Interns Gain Equity While Building Cutting‑Edge AI

Kimi, a three‑year‑old AI‑native unicorn valued over $120 billion, launches a “Time‑Machine” option program that grants interns equity while showcasing its rapid valuation growth, record‑breaking context lengths, novel Kimi Linear architecture, token‑efficiency gains, and open‑source models that rival leading LLMs.

AI Talent ProgramAgent SwarmsAttention Residuals
0 likes · 10 min read
Kimi’s ‘Option Time Machine’: Interns Gain Equity While Building Cutting‑Edge AI
SuanNi
SuanNi
Mar 17, 2026 · Artificial Intelligence

How Attention Residuals Boost Transformer Efficiency and Scale

The article presents the Attention Residuals architecture, explains how it replaces uniform residual addition with learned attention‑based aggregation, details full and block variants, engineering tricks for distributed training, and shows extensive scaling‑law experiments where the new design consistently improves validation loss and training efficiency across model sizes.

Attention ResidualsDeep LearningEfficient Training
0 likes · 13 min read
How Attention Residuals Boost Transformer Efficiency and Scale
ShiZhen AI
ShiZhen AI
Mar 17, 2026 · Artificial Intelligence

Kimi’s Attention Residuals Swap a Decade-Old Residual Trick for 1.25× Faster 48B MoE

The Kimi team introduces Attention Residuals, a softmax‑based replacement for the uniform residual connections used in Transformers for a decade, enabling selective aggregation of layer histories, reducing hidden‑state growth, and achieving a 1.25× compute‑efficiency gain on a 48‑billion‑parameter MoE model with less than 2% inference latency increase.

Attention ResidualsCompute EfficiencyDeep Learning
0 likes · 10 min read
Kimi’s Attention Residuals Swap a Decade-Old Residual Trick for 1.25× Faster 48B MoE