Tagged articles

BarunLM-35M

2 articles · Page 1 of 1
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Aug 2, 2026 · Artificial Intelligence

A 35M-Parameter Model Trained on a Single GPU Claims Best Sub-100M Performance

Developer Harshal Singh released BarunLM-35M, a 35-million-parameter language model that fits on an ESP32-S3, achieves 41.01% average accuracy on nine zero-shot benchmarks—outperforming larger 160-M-parameter models—using a single H200 GPU, with novel alternating local/global attention and a learnable residual selector.

BarunLM-35MGPU traininglocal attention
0 likes · 6 min read
A 35M-Parameter Model Trained on a Single GPU Claims Best Sub-100M Performance
Machine Heart
Machine Heart
Aug 2, 2026 · Artificial Intelligence

Can a Single GPU Yield the Best Sub‑100M Parameter Language Model?

BarunLM‑35M, a 35‑million‑parameter language model trained on a single H200 GPU, achieves a 41.01% average score on nine zero‑shot benchmarks, surpassing much larger models, thanks to a hybrid local‑global attention scheme, learnable residual selector, and other architectural optimizations.

AI researchBarunLM-35Mlocal-global attention
0 likes · 6 min read
Can a Single GPU Yield the Best Sub‑100M Parameter Language Model?