Machine Heart
Aug 2, 2026 · Artificial Intelligence
Can a Single GPU Yield the Best Sub‑100M Parameter Language Model?
BarunLM‑35M, a 35‑million‑parameter language model trained on a single H200 GPU, achieves a 41.01% average score on nine zero‑shot benchmarks, surpassing much larger models, thanks to a hybrid local‑global attention scheme, learnable residual selector, and other architectural optimizations.
AI researchBarunLM-35Mlocal-global attention
0 likes · 6 min read
