Tagged articles

sliding window attention

2 articles · Page 1 of 1
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Sep 4, 2026 · Artificial Intelligence

Sliding Window Attention Beats Linear Attention: Microsoft's 0-Training Baseline Matches Costly Post-Training

Microsoft research shows sliding window attention with four attention sinks achieves 99% performance recovery on knowledge tasks without any post-training, matching linear attention methods requiring hundreds of millions of tokens while outperforming them 2-10x on long-context benchmarks.

KV CacheLinear AttentionMicrosoft research
0 likes · 9 min read
Sliding Window Attention Beats Linear Attention: Microsoft's 0-Training Baseline Matches Costly Post-Training
Baobao Algorithm Notes
Baobao Algorithm Notes
Jul 31, 2024 · Artificial Intelligence

What Makes Mistral’s 7B, Mixtral, and Large 2 Models Stand Out? A Deep Technical Dive

This article compiles key technical details of the Mistral model family—including Mistral 7B, Mixtral 8×7B, Mixtral 8×22B, Mistral Nemo, and Mistral Large 2—covering their architectural innovations such as sliding‑window attention, grouped‑query attention, mixture‑of‑experts design, scaling parameters, performance benchmarks, quantization requirements, and practical deployment commands.

Grouped Query AttentionMistralMixtral
0 likes · 17 min read
What Makes Mistral’s 7B, Mixtral, and Large 2 Models Stand Out? A Deep Technical Dive