AI Engineering
Aug 23, 2026 · Artificial Intelligence
FreeToken Runs a 35B MoE Model on an 8 GB GPU at 39 tokens/s – Is Local Freedom Real?
FreeToken, an open‑source MoE inference engine from UC Berkeley, enables a 35B model to run on an 8 GB GPU at 39.3 tokens per second by offloading weights to system memory and dynamically scheduling between CPU and GPU, offering a detailed performance analysis and community feedback.
CPU OffloadFreeTokenGPU
0 likes · 6 min read
