Tagged articles

KV Cache Eviction

1 articles · Page 1 of 1
Machine Heart
Machine Heart
Sep 19, 2026 · Artificial Intelligence

Random Attention: Random KV Cache Eviction Rivals Best Baselines, Boosts Throughput 43%

Salesforce and UIUC researchers propose Random Attention, a simple KV cache eviction method that randomly retains reasoning tokens while fully protecting the prompt, matching the accuracy of complex importance-based methods across multiple models and tasks while increasing inference throughput by 32-43% in vLLM serving.

Attention MechanismCache CompressionEfficient AI
0 likes · 18 min read
Random Attention: Random KV Cache Eviction Rivals Best Baselines, Boosts Throughput 43%