How DeepSeek mHC Uses Its Residual Streams: Selective Routing and Near-Identity Mixing

Analysis of DeepSeek-V4-Flash's mHC architecture shows its four residual streams are not uniformly utilized; read/write weights concentrate on ~2 streams per layer, and deep-layer residual mixing matrices approach identity. Interventions removing weakest paths or replacing deep mixing with identity cause minimal performance drops (≤0.38% avg score), suggesting simplification opportunities.

Machine Heart
Machine Heart
Machine Heart
How DeepSeek mHC Uses Its Residual Streams: Selective Routing and Near-Identity Mixing

Paper and Model

The study analyzes DeepSeek-V4-Flash-0731 , which employs Manifold-Constrained Hyper-Connections (mHC) with four parallel residual streams. The paper How Does mHC Use Its Residual Streams? Selective Routing and Near-Identity Mixing (arXiv:2609.05309v1) investigates how these streams are actually used during inference.

Methodology

Using 512 sequences of length 1,024 from the C4 dataset, the authors record per-token read/write weights and residual mixing matrices for every Attention and FFN sub-layer. Performance is measured via C4 perplexity (PPL) and six zero-shot tasks: ARC-Easy, ARC-Challenge, PIQA, HellaSwag, MMLU, and GSM8K.

Read/Write Routing Concentration

For each sub-layer, the dominant stream (highest read/write weight) is highly consistent across tokens: average consistency 0.871 for read, 0.905 for write. The effective number of streams (entropy-based) averages 1.998 for read and 1.775 for write, meaning a typical sub-layer concentrates weight on roughly two of the four streams. Figure 2 illustrates these metrics across depth.

Figure 2: Dominant stream consistency and effective stream count per sub-layer
Figure 2: Dominant stream consistency and effective stream count per sub-layer

Stream Representation Similarity

Cosine similarity between the four streams' hidden states starts at 1 (identical initialization) then diverges after the first layer. Across the network, the six pairwise similarities average 0.404 (range 0.235–0.564), indicating streams develop distinct directional representations and do not reconverge. Figure 4 shows the per-pair evolution.

Figure 4: Pairwise cosine similarity of residual streams across depth
Figure 4: Pairwise cosine similarity of residual streams across depth

Residual Mixing Matrices

Deviation from identity matrix is measured by mean absolute difference. Shallow layers (0–21) show notable off-diagonal weights (cross-stream mixing). From layer 22 onward, matrices approach identity: in layers 23–42, 32 of 40 sub-layers have deviation <0.01. Figure 5 plots deviation per sub-layer; Figure 6 contrasts average matrices for shallow vs. deep layers.

Figure 5: Identity deviation per Attention and FFN sub-layer
Figure 5: Identity deviation per Attention and FFN sub-layer
Figure 6: Average residual mixing matrices for layers 0–21 and 22–42
Figure 6: Average residual mixing matrices for layers 0–21 and 22–42

Intervention Experiments

Read/Write Path Pruning

For each token and sub-layer, only the top- k read/write paths are kept and renormalized. Top-3 pruning has negligible impact. Top-2 pruning increases PPL and reduces average task score more noticeably, with write pruning having larger effect. The weakest path can be approximately removed.

Figure: PPL and task score changes under top-k pruning
Figure: PPL and task score changes under top-k pruning

Residual Mixing Replacement

Deep layers (22–42) → identity: PPL +1.9%, average task score +0.04 pp.

All layers → identity: PPL +42.3%, average task score −3.28 pp.

Shallow layers fixed to per-sub-layer C4 average matrix (removing per-token dynamics): PPL +0.2%, average task score −0.25 pp.

Figure 7 details per-task score changes for each intervention.

Figure 7: Zero-shot score changes per task for each intervention
Figure 7: Zero-shot score changes per task for each intervention

Discussion and Simplification Directions

The results suggest two simplification avenues: (1) reduce read/write routing density while retaining multi-stream structure, and (2) make residual mixing dynamic only in shallow layers, using identity in deep layers. Recent models already adopt simpler residual pathways: Tencent Hy4-preview's identity Hyper-Connections (iHC) fixes the mixing matrix to identity, and Qwen3.8-Flash-Next's Gated Residual removes it entirely—both equivalent to setting the mixing matrix to identity in this paper's notation.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

model analysisarchitecture simplificationHyper-ConnectionsmHCDeepSeek-V4-Flashresidual streamsidentity mixingselective routing
Machine Heart
Written by

Machine Heart

Professional AI media and industry service platform

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.