Sliding Window Attention Beats Linear Attention: Microsoft's 0-Training Baseline Matches Costly Post-Training
Microsoft research shows sliding window attention with four attention sinks achieves 99% performance recovery on knowledge tasks without any post-training, matching linear attention methods requiring hundreds of millions of tokens while outperforming them 2-10x on long-context benchmarks.
