Machine Heart
Jul 20, 2026 · Artificial Intelligence
HiLS-Attention: Mathematically Correct Sparse Attention with 13‑15× Faster Inference and 512× Extrapolation
HiLS-Attention introduces a hierarchical landmark sparse attention that mathematically resolves chunk‑importance estimation and end‑to‑end differentiability, matching full‑attention perplexity while delivering 13.5× faster prefill, 15.7× faster decode, and up to 512‑fold context extrapolation, even surpassing full attention on long‑context retrieval tasks.
EfficiencyHiLS-Attentionlanguage modeling
0 likes · 11 min read
