Tagged articles

streaming weight loading

1 articles · Page 1 of 1
Machine Heart
Machine Heart
Aug 1, 2026 · Artificial Intelligence

How a 128 GB Mac Loaded 1.6 TB of Kimi K3 Weights and Ran Inference

A developer demonstrated that a 128 GB M5 Max Mac can stream‑load the 1.6 TB MXFP4 weights of the 2.8‑trillion‑parameter Kimi K3 model, achieving 0.32 token/s, while an 80‑GPU RTX 5090 cluster reaches 20 token/s, highlighting both feasibility and speed limits of large‑model inference on consumer hardware.

GPU clusterKimi K3Mac M5 Max
0 likes · 6 min read
How a 128 GB Mac Loaded 1.6 TB of Kimi K3 Weights and Ran Inference