Analysis: DeepSeek-V4-Flash's mHC Actually Uses Only ~2 of Its 4 Residual Streams
机器之心 · wechat · 2026-09-27
A new paper, How Does mHC Use Its Residual Streams?, dissects the multi-stream residual architecture (mHC) in DeepSeek-V4-Flash and finds it substantially over-provisioned.
Key findings
- Read/write routing concentrates on 2 of 4 residual streams (effective streams: 1.998 read / 1.775 write), with 0.87–0.91 consistency in dominant-stream choice per sublayer.
- Residual mixing shows strong depth dependence: layers 22–42 are near-identity (32 of 40 sublayers deviate <0.01), while shallow layers retain cross-stream mixing.
- Stream representations diverge early and stay distinct (mean pairwise cosine similarity 0.404).
Interventions
- Pruning the weakest read/write path: PPL up at most 2.7%, six-task average down ≤0.38 points.
- Replacing deep-layer mixing matrices with the identity matrix: PPL +1.9%, scores essentially unchanged (84.22→84.26); replacing all layers costs 42.3% PPL.
- Fixing shallow layers to their C4 token-average matrix: only +0.2% PPL, showing per-token dynamic mixing adds little.
Implication: routing density and mixing dynamics could be simplified depth-dependently. Tencent's Hy4-preview (identity Hyper-Connections) and Qwen3.8-Flash-Next's Gated Residual already move in this direction.
More from Models
- TeleOCR Trends on Hugging Face: A Qwen2.5-VL-Based Chinese Document OCR Model — XingChen-AGI · 2026-09-28
- Kaggle Game Arena: Evaluating LLMs via Head-to-Head Chess, Poker, and Werewolf — kaggle · 2026-09-28
- Perplexity CEO: still using sol 6 for knowledge work — cheap, fast, great compaction — gabriel1 · 2026-09-28
- NerfBench's First Results Find No Nerf: Claude Opus 5.5 Dips Just 0.8% vs Launch — alejandroll10 · 2026-09-28
- Most humans can read this image instantly — most AI vision models can't — JeremyNguyenPhD · 2026-09-28
- AI-generated 7-minute SQLite repo explainer stuns with coherent code walkthrough — deedydas · 2026-09-28