Machine Learning Algorithms & Natural Language Processing
Aug 2, 2026 · Artificial Intelligence
A 35M-Parameter Model Trained on a Single GPU Claims Best Sub-100M Performance
Developer Harshal Singh released BarunLM-35M, a 35-million-parameter language model that fits on an ESP32-S3, achieves 41.01% average accuracy on nine zero-shot benchmarks—outperforming larger 160-M-parameter models—using a single H200 GPU, with novel alternating local/global attention and a learnable residual selector.
BarunLM-35MGPU traininglocal attention
0 likes · 6 min read
