Tagged articles

NoPE

3 articles · Page 1 of 1
AI Engineering
AI Engineering
Aug 2, 2026 · Artificial Intelligence

Kimi K3 Technical Report Reveals Answers to Key Architecture Questions

The 47‑page Kimi K3 technical report, released on July 27, details the 2.8‑trillion‑parameter model’s novel LatentMoE, SiTU‑GLU, Quantile Balancing, Attention Residuals, and full‑stack NoPE design, explains how these solve activation‑explosion and load‑imbalance problems, and provides open‑source code for inference and agentic RL.

Attention ResidualsKimi K3LatentMoE
0 likes · 12 min read
Kimi K3 Technical Report Reveals Answers to Key Architecture Questions
Data Party THU
Data Party THU
Aug 11, 2025 · Artificial Intelligence

What Sets the Latest LLMs Apart? A Deep Dive into V3, OLMo, Gemma, Mistral, Llama 4 and More

This article systematically compares the architectures of recent large language models—including DeepSeek V3/R1, OLMo 2, Gemma 3, Mistral Small 3.1, Llama 4, Qwen 3, SmolLM 3 and Kimi 2—highlighting innovations such as MLA, MoE, post‑norm, sliding‑window attention, NoPE and optimizer choices, with diagrams and code examples to illustrate their impact on efficiency and performance.

ComparisonLLMMLA
0 likes · 12 min read
What Sets the Latest LLMs Apart? A Deep Dive into V3, OLMo, Gemma, Mistral, Llama 4 and More
DevOps
DevOps
Apr 7, 2025 · Artificial Intelligence

Meta Llama 4 Scout, Maverick, and Behemoth: Architecture, NoPE Innovation, and Training Advances

The article introduces Meta's newly open‑sourced Llama 4 series—including Scout with a 1 billion‑token context window, Maverick with 400 billion parameters, and the upcoming Behemoth teacher model—detailing their expert‑mix architecture, the NoPE positional‑encoding removal, training pipelines, performance benchmarks, and infrastructure improvements for large‑scale AI research.

AI researchLlama 4NoPE
0 likes · 8 min read
Meta Llama 4 Scout, Maverick, and Behemoth: Architecture, NoPE Innovation, and Training Advances