Machine Learning Algorithms & Natural Language Processing
Sep 21, 2026 · Artificial Intelligence
Random KV Cache Eviction Rivals Top Baselines, Boosts Throughput 43%
Salesforce AI Research and UIUC propose Random Attention, a simple KV cache eviction method that randomly retains reasoning tokens while protecting the prompt, matching state-of-the-art baselines across math, science, and code reasoning tasks and improving inference throughput by 32–43% in vLLM serving.
Cache EvictionKV CacheLLM Inference
0 likes · 20 min read
