Machine Heart
Sep 19, 2026 · Artificial Intelligence
Random Attention: Random KV Cache Eviction Rivals Best Baselines, Boosts Throughput 43%
Salesforce and UIUC researchers propose Random Attention, a simple KV cache eviction method that randomly retains reasoning tokens while fully protecting the prompt, matching the accuracy of complex importance-based methods across multiple models and tasks while increasing inference throughput by 32-43% in vLLM serving.
Attention MechanismCache CompressionEfficient AI
0 likes · 18 min read
