Hybrid Attention: Why Kimi and DeepSeek Now Share a Model Architecture
This article traces the evolution of attention mechanisms in large language models, showing how hybrid architectures now combine linear and sparse attention — exemplified by GLM-5.3-Flash integrating Kimi's KDA and DeepSeek's DSA — driven by shifting constraints from context length to agent workloads, with MiniMax's architectural journey illustrating the trade-offs.
