Tagged articles

activated parameters

1 articles · Page 1 of 1
Cambridge Mofang Notes
Cambridge Mofang Notes
Sep 1, 2026 · Artificial Intelligence

From Dense to MoE: Decoding Total vs. Activated Parameters

This article explains the distinction between total and activated parameters in Mixture-of-Experts (MoE) models, contrasting dense and sparse architectures, detailing expert routing mechanisms, and analyzing memory and compute implications across model loading, prefill, and decode stages.

Mixture of ExpertsMoEactivated parameters
0 likes · 15 min read
From Dense to MoE: Decoding Total vs. Activated Parameters