How DeepSeek mHC Uses Its Residual Streams: Selective Routing and Near-Identity Mixing
Analysis of DeepSeek-V4-Flash's mHC architecture shows its four residual streams are not uniformly utilized; read/write weights concentrate on ~2 streams per layer, and deep-layer residual mixing matrices approach identity. Interventions removing weakest paths or replacing deep mixing with identity cause minimal performance drops (≤0.38% avg score), suggesting simplification opportunities.
Paper and Model
The study analyzes DeepSeek-V4-Flash-0731 , which employs Manifold-Constrained Hyper-Connections (mHC) with four parallel residual streams. The paper How Does mHC Use Its Residual Streams? Selective Routing and Near-Identity Mixing (arXiv:2609.05309v1) investigates how these streams are actually used during inference.
Methodology
Using 512 sequences of length 1,024 from the C4 dataset, the authors record per-token read/write weights and residual mixing matrices for every Attention and FFN sub-layer. Performance is measured via C4 perplexity (PPL) and six zero-shot tasks: ARC-Easy, ARC-Challenge, PIQA, HellaSwag, MMLU, and GSM8K.
Read/Write Routing Concentration
For each sub-layer, the dominant stream (highest read/write weight) is highly consistent across tokens: average consistency 0.871 for read, 0.905 for write. The effective number of streams (entropy-based) averages 1.998 for read and 1.775 for write, meaning a typical sub-layer concentrates weight on roughly two of the four streams. Figure 2 illustrates these metrics across depth.
Stream Representation Similarity
Cosine similarity between the four streams' hidden states starts at 1 (identical initialization) then diverges after the first layer. Across the network, the six pairwise similarities average 0.404 (range 0.235–0.564), indicating streams develop distinct directional representations and do not reconverge. Figure 4 shows the per-pair evolution.
Residual Mixing Matrices
Deviation from identity matrix is measured by mean absolute difference. Shallow layers (0–21) show notable off-diagonal weights (cross-stream mixing). From layer 22 onward, matrices approach identity: in layers 23–42, 32 of 40 sub-layers have deviation <0.01. Figure 5 plots deviation per sub-layer; Figure 6 contrasts average matrices for shallow vs. deep layers.
Intervention Experiments
Read/Write Path Pruning
For each token and sub-layer, only the top- k read/write paths are kept and renormalized. Top-3 pruning has negligible impact. Top-2 pruning increases PPL and reduces average task score more noticeably, with write pruning having larger effect. The weakest path can be approximately removed.
Residual Mixing Replacement
Deep layers (22–42) → identity: PPL +1.9%, average task score +0.04 pp.
All layers → identity: PPL +42.3%, average task score −3.28 pp.
Shallow layers fixed to per-sub-layer C4 average matrix (removing per-token dynamics): PPL +0.2%, average task score −0.25 pp.
Figure 7 details per-task score changes for each intervention.
Discussion and Simplification Directions
The results suggest two simplification avenues: (1) reduce read/write routing density while retaining multi-stream structure, and (2) make residual mixing dynamic only in shallow layers, using identity in deep layers. Recent models already adopt simpler residual pathways: Tencent Hy4-preview's identity Hyper-Connections (iHC) fixes the mixing matrix to identity, and Qwen3.8-Flash-Next's Gated Residual removes it entirely—both equivalent to setting the mixing matrix to identity in this paper's notation.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
