Tagged articles

Cache Compression

2 articles · Page 1 of 1
Machine Heart
Machine Heart
Sep 19, 2026 · Artificial Intelligence

Random Attention: Random KV Cache Eviction Rivals Best Baselines, Boosts Throughput 43%

Salesforce and UIUC researchers propose Random Attention, a simple KV cache eviction method that randomly retains reasoning tokens while fully protecting the prompt, matching the accuracy of complex importance-based methods across multiple models and tasks while increasing inference throughput by 32-43% in vLLM serving.

Attention MechanismCache CompressionEfficient AI
0 likes · 18 min read
Random Attention: Random KV Cache Eviction Rivals Best Baselines, Boosts Throughput 43%
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Aug 18, 2026 · Artificial Intelligence

Why DeepSeek’s Cache Costs Jumped 11‑Fold: Long‑Context Surge and the New “Storage Tax”

DeepSeek raised its cache‑hit price up to 11 times as exploding long‑context demand forces a shift to tiered KV storage, exposing hidden storage, I/O and scheduling costs that turn GPU compute into costly data‑movement, prompting developers to rethink cache strategies.

Cache CompressionDeepSeekHot-Cold Tiering
0 likes · 10 min read
Why DeepSeek’s Cache Costs Jumped 11‑Fold: Long‑Context Surge and the New “Storage Tax”