Machine Learning Algorithms & Natural Language Processing
Aug 29, 2026 · Artificial Intelligence
Dropping Intermediate Tokens: How Prefix Sliding Achieves Up to 3× Faster Long-Context Reasoning
Prefix Sliding keeps the task prefix and a sliding window of recent tokens while evicting older intermediate tokens from the KV cache, enabling up to three‑fold speedups for long‑chain inference without retraining and extending reinforcement‑learning rollouts beyond 100 k tokens.
AttentionKV cacheLong-context inference
0 likes · 11 min read
