Run 744B MoE Model on 25GB RAM: Colibri's Tiered Storage Breakthrough
Colibri, a pure C inference engine with zero dependencies, enables running the 744B parameter GLM-5.2 MoE model on consumer hardware with just 25GB RAM and NVMe SSD by leveraging MoE sparsity and a three-tier storage scheduling system across VRAM, RAM, and disk, achieving 0.05–6.8 tok/s depending on hardware.
