Artificial Intelligence
AI FinOps 2.0
48 min read

Running 744B MoE on 25GB RAM: Int4, Disk Streaming & MLA Compression

This article details how a custom C inference engine runs the 744B-parameter GLM 5.2 MoE model on 25GB RAM by quantizing routed experts to int4, streaming them from disk via pread, compressing KV cache 57x with Multi-head Latent Attention, and employing a five-tier memory hierarchy with speculative decoding.

Data STUDIO
Data STUDIO
Data STUDIO
Running 744B MoE on 25GB RAM: Int4, Disk Streaming & MLA Compression

01 Why 744B Fits in Small Memory

GLM 5.2 is a 744B-parameter Mixture-of-Experts (MoE) model with 78 layers (first 3 dense, remaining 75 with 256 routed experts each). Each token activates only top-8 experts per layer, so only ~17B parameters (≈9.9 GB at int4) are always resident; the remaining ~727B expert parameters (≈362 GB) stay on disk. Per-token expert working set is ~11 GB, with strong routing locality enabling caching.

An expert comprises gate, up, down matrices (shape O×I). At int4 (2 values/byte) plus per-row scales, one expert ≈19 MB — the unit of cold reads and cache granularity.

02 Model & Hardware

Minimum viable hardware: any modern x86-64 (AVX2) or Apple Silicon, 16–26 GB RAM, NVMe SSD, no GPU required. Benchmarks: Framework 13 laptop ~0.37 tok/s; desktop ~0.1–0.3 tok/s. Development used a 4×NVIDIA L40, AMD EPYC 7763 (124 cores, AVX2, no AVX-512), 228 GB RAM, 3.2 TB NVMe. Software: GCC 13.3, NVCC 12.8, Python 3.12.

Key model config: hidden_size=6144, 78 layers, 256 experts/layer, top-8, 1 shared expert, MLA with kv_lora_rank=512, qk_nope_head_dim=192, qk_rope_head_dim=64, vocab=154880, FP8 (e4m3) checkpoint.

03 Single-File C Inference Engine

Core engine is ~3900 lines in glm.c plus header-only helpers ( st.h, tier.h), no BLAS or frameworks. Builds to a 377 KB static binary. GPU build links backend_cuda.cu; CPU build uses OpenMP. Self-tests verify JSON parsing, safetensors primitives, tier logic, grammar, decode batch, and AVX2 int8 dot-product exactness (bitwise match with reference).

OpenMP tuning: default passive wait hurts tiny expert matmuls; engine re-execs with active spin, reducing matmul time from 66.9 s to 20.9 s.

04 Handwritten Math Kernels

RMSNorm, Softmax, SiLU, and RoPE implemented in C. Transformer layer sequence: input norm → attention → residual → post norm → MoE/dense FFN → residual. RoPE rotates query/key pairs by position-dependent angles.

05 Streaming Validation with 400-Line olmoe.c

Before full engine, a minimal streaming engine validated the design: resident dense weights, expert cache (quantized gate/up/down + per-row scales), LRU eviction, on-demand pread from disk. Quantization uses per-row absmax to int8. MoE step: router logits → softmax → top-K → fetch experts → SwiGLU → weighted sum. Test on small model: 20/20 tokens matched PyTorch oracle, confirming streaming load, cache, and decode correctness.

06 FP8 → Int4 Conversion

Vendor checkpoint is FP8 (e4m3, block size 128×128) ≈756 GB. Offline Python converter decodes FP8 blocks to float32, then per-row symmetric quantizes to int4 ([-8,7]), packs two nibbles/byte (mapped to 0..15). Precision allocation: routed experts → int4; embeddings, output head, MTP → int8; router, norms, biases → float32. Conversion runs per shard to avoid holding both FP8 and int4. Self-test: mean relative error 0.0226. Final layout: ~362 GB int4 experts, ~12 GB int8 I/O & MTP, <1 GB float32 router/norm/bias.

07 Int4 Trade-offs

Teacher-forcing alignment on small model (32 positions): experts@16/dense@16 → 32/32; experts@8/dense@8 → 30/32; experts@4/dense@8 → 28/32; experts@4/dense@4 → 9/32; experts@8/dense@4 → 6/32. Routed experts tolerate int4 better; dense path (attention, shared experts) is more sensitive. Python converter ( np.rint) and C runtime ( lrintf) produce identical int4 bytes (9/32 match).

08 On-Demand Weight Reading

144 shards, ~120K tensors. Startup builds hash index ( st_find). Runtime reads only required tensor bytes via pread; posix_fadvise(..., POSIX_FADV_DONTNEED) drops pages after use. O_DIRECT bypasses page cache: 8-thread cold read 64×19 MB blocks at 4.13 GB/s (4.8 ms/block) vs buffered 3.22 GB/s (6.2 ms/block). 8 experts × 75 layers ≈600 reads per token; tiering aims to avoid cold reads.

09 MLA: KV Cache Compressed 57×

Multi-head Latent Attention stores per-token latent Lc (512-d) + rotary Rc (64-d) = 576 floats vs full K/V 64×(256+256)=32768 floats. Decode uses weight absorption: key up-projection absorbed into query, scoring directly against latent. Verified: absorbed and explicit K/V reconstruction produce identical token sequences. KV cache persists to disk ( .glm_kv), enabling cross-process resume (57 tokens restored in 0.0 s).

10 Sparse Attention Indexer

For long contexts, indexer scores past keys via per-head ReLU'd dot products, keeps top index_topk=2048. Below 2048 tokens, equivalent to dense attention. Toggle tests: DSA_FORCE=1 and DSA=0 produce identical generated text.

11 MoE Routing & Expert Execution

Router: sigmoid(logit) + correction_bias for top-K selection; weighting uses plain sigmoid(logit) scaled by routed_scaling_factor=2.5. Batch prefill/verification: unique union of experts across batch, each loaded once. Expert = SwiGLU (gate·SiLU(gate) × up → down). Routing statistics show heavy skew (hottest expert triggered 776 times), validating cache effectiveness.

12 AVX2 Integer Kernels

Matmul dispatcher: GPU → int8/int4 AVX2 → float fallback. Activation quantized per-row to int8 (absmax/127). AVX2 lacks signed byte dot-product; uses sign-folding:

_mm256_maddubs_epi16(_mm256_sign_epi8(w,w), _mm256_sign_epi8(x,w))

then horizontal sum. test_idot passes (bitwise exact). Int8 activation path adds ~0.3% RMS noise, occasionally changing sampled tokens (speed vs exact reproducibility trade-off).

13 Five-Tier Expert Residency

Tiers: VRAM → pinned RAM → per-layer LRU cache → OS page cache → disk. Swap policy: frequency-primary, recency tie-breaker (LFU+recency score = (heat<<8) | recent). Hot expert must exceed cold by >25%+4 margin. Contiguous expert matrices (gate/up/down) read in single ~19 MB pread into slab, three QT structs point into it. Memory budget explicitly reserves dense weights, KV cache, working set, page cache reserve (2.5 GB), activation slack (1.2 GB). Example: 20 GB budget → cap 1 expert/layer (projected peak 19.7 GB); 200 GB budget → cap 45/layer (peak 199.3 GB). Four L40 placement: 185 GB VRAM hot tier (9780/15860 experts), 183 GB RAM warm experts, 10.9 GB dense, 6.1 GB runtime. Startup doctor checks viability; usage history ( .glm_usage) pre-pins hot experts at startup.

14 Handling Disk Latency

Page cache is primary latency buffer: cold run 24 tokens 76.2 s (expert-disk 43.7 s); warm repeat 50.2 s (expert-disk 21.9 s); forced DONTNEED 184.8 s (expert-disk 149.5 s). Prefetch/I/O worker/lookahead tested: on this virtualized NVMe, pipelined load slowed throughput (0.38 → 0.16 tok/s); lookahead raised hit rate 58%→62% but eviction pressure lowered throughput. Best single-thread I/O mode: direct 0.23 tok/s, buffered 0.17, drop 0.20, mmap 0.19. Results are device-specific; cannot extrapolate.

15 Three Speculative Decoding Methods

Draft sources: model's MTP head (must stay int8, else acceptance ~0%), n-gram, grammar (GBNF). Draft tokens verified in one batched forward; longest matching prefix accepted. MTP ON (draft=3): 2.67 tokens/forward, 50% acceptance, 0.27 tok/s; MTP OFF: 1.07, 0%, 0.25 tok/s. Grammar example: 50% acceptance (3/6 forced drafts). Sampling temperature critical: greedy (temp=0) stable; temp≥1.3 degrades into gibberish, attributed to int4 quantization noise in distribution tails.

16 OpenAI-Compatible Server

Python stdlib HTTP server handles scheduling, queueing, SSE streaming, backpressure. C engine runs as persistent subprocess, communicates via line protocol ( SUBMIT id slot bytes max_tokens temperature top_p). Slot = single mutable KV context; fixed capacity, bounded queue, HTTP 429 when full. Logs show prefill layers, then steady streaming. Example request: 28-token prompt, TTFT 5.68 s, 200 completion tokens, mean inter-token 0.496 s, decode 2.02 tok/s. Concurrency scaling sublinear: 1 req 2.08 tok/s, 4 reqs 2.9 tok/s total (1.39×), latency 72→206 s; 20 concurrent → 5 admitted, 15 HTTP 429.

17 Real Conversations & Performance Boundaries

Sample queries: Python prime function (119 s, 200 tokens), train speed unit conversion (correct), Australian capital (3-sentence limit obeyed), haiku (5-7-5). Persistent KV recall verified (hummingbird fact). Smoke test: 3 questions, 66.7% accuracy (not a full benchmark). Same hardware, three placements: CPU 20 GB (0.28 tok/s, hit 3.5%, bottleneck disk 46.5 s); CPU 200 GB (0.28 tok/s, hit 68.6%, disk 36 s, matmul 36 s); GPU 4×L40 (kernel 289.7 ms, H2D 36.6 ms, D2H 58.2 ms, bottleneck GPU compute). Thread scaling peaks at ~32 threads (IPC 2.11→0.23), classic memory wall. Other devices: Framework 13 0.37 tok/s, Ryzen 9950X 0.10–0.28, i5-12600K 0.08, Apple M5 Max/Metal 2.06 tok/s (uncontrolled conditions).

18 Correctness Verification

Tiny reference model (same architecture, smaller weights) in PyTorch. Key tensors: q_a_proj, kv_a_proj_with_mqa, kv_b_proj, indexer.wq_b, experts.gate_up_proj, gate.e_score_correction_bias, shared_experts.gate_proj. Prefill teacher-forcing: 32/32 positions matched. Greedy decode: 20/20 tokens matched (expert cache hit 88.1%, 227.5 tok/s). Notes: loading-time banners may differ; int8 activation path may cause single-token divergence due to quantization noise and FP accumulation order.

19 From Checkpoint to Service

Pipeline: 1) make glm; 2) tiny oracle test SNAP=./glm_tiny TF=1 ./glm 64 16 16; 3) convert FP8→int4 per shard

python tools/convert_fp8_to_int4.py --repo zai-org/GLM-5.2-FP8 --outdir /nvme/glm52_i4 --ebits 4 --io-bits 8

(plus MTP); 4) resource plan & doctor check; 5) run 20 GB CPU-only: PROMPT=... NGEN=24 RAM_GB=20 SNAP=/nvme/glm52_i4 ./glm 64; 6) GPU tier:

GLM_CUDA=1 GLM_GPUS=0,1,2,3 CUDA_EXPERT_GB=185 RAM_GB=200 ...

; 7) OpenAI server:

GLM_MODEL=/nvme/glm52_i4 GLM_API_KEY=local-secret python openai_server.py --host 127.0.0.1 --port 8000 --kv-slots 4

.

Key insight: Not "stuffing 744B into 20 GB" but a memory-hierarchy engineering: int4 shrinks experts, disk holds sleeping experts, pread +cache manages working set, MLA compresses context state, RAM/VRAM split hot experts, GPU/CPU does matmul. 20 GB mode runs but bottlenecks on disk cold reads; more RAM shifts bottleneck to CPU matmul; more VRAM moves expert compute to GPU. Four L40s are dev/measurement hardware, not minimum requirement.

C implementation — GitHub

Model checkpoint — Hugging Face

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

speculative decodingMixture of ExpertsMLAKV cache compressionint4 quantizationGLM-5.2disk streamingAVX2 kernelsC inference engine
Data STUDIO
Written by

Data STUDIO

Click to receive the "Python Study Handbook"; reply "benefit" in the chat to get it. Data STUDIO focuses on original data science articles, centered on Python, covering machine learning, data analysis, visualization, MySQL and other practical knowledge and project case studies.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.