Tagged articles

RMSNorm

3 articles · Page 1 of 1
ThinkingAgent
ThinkingAgent
Aug 4, 2026 · Artificial Intelligence

How Transformers Compute Contextual Relationships

The article explains how the Transformer architecture replaces RNNs with self‑attention, detailing the Q‑K‑V mechanism, positional encodings such as RoPE, multi‑head attention, modern improvements like SwiGLU and RMSNorm, and provides formulas for parameter and FLOP estimation.

FFNMulti-Head AttentionPositional Encoding
0 likes · 26 min read
How Transformers Compute Contextual Relationships
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Mar 20, 2026 · Artificial Intelligence

Why Kimi Dropped Residual Connections: A First‑Person Deep Dive into Attention Residuals

This article explains how Attention Residuals (AttnRes) replace traditional residual shortcuts with layer‑wise attention, details the mathematical reformulation, design constraints, static‑Q trick, full and block variants, and presents experimental evidence of significant accuracy gains with modest overhead.

AttentionNLPRMSNorm
0 likes · 11 min read
Why Kimi Dropped Residual Connections: A First‑Person Deep Dive into Attention Residuals