Tagged articles

HiLS-Attention

1 articles · Page 1 of 1
Machine Heart
Machine Heart
Jul 20, 2026 · Artificial Intelligence

HiLS-Attention: Mathematically Correct Sparse Attention with 13‑15× Faster Inference and 512× Extrapolation

HiLS-Attention introduces a hierarchical landmark sparse attention that mathematically resolves chunk‑importance estimation and end‑to‑end differentiability, matching full‑attention perplexity while delivering 13.5× faster prefill, 15.7× faster decode, and up to 512‑fold context extrapolation, even surpassing full attention on long‑context retrieval tasks.

EfficiencyHiLS-Attentionlanguage modeling
0 likes · 11 min read
HiLS-Attention: Mathematically Correct Sparse Attention with 13‑15× Faster Inference and 512× Extrapolation