Data STUDIO
Sep 11, 2026 · Artificial Intelligence
Running 744B MoE on 25GB RAM: Int4, Disk Streaming & MLA Compression
This article details how a custom C inference engine runs the 744B-parameter GLM 5.2 MoE model on 25GB RAM by quantizing routed experts to int4, streaming them from disk via pread, compressing KV cache 57x with Multi-head Latent Attention, and employing a five-tier memory hierarchy with speculative decoding.
AVX2 kernelsC inference engineGLM-5.2
0 likes · 48 min read
