Random Attention: Random KV Cache Eviction Rivals Best Baselines, Boosts Throughput 43%
Salesforce and UIUC researchers propose Random Attention, a simple KV cache eviction method that randomly retains reasoning tokens while fully protecting the prompt, matching the accuracy of complex importance-based methods across multiple models and tasks while increasing inference throughput by 32-43% in vLLM serving.
