Ling-3.0 Architecture Revealed: Native Hybrid Attention with 256K Context
SonglinYang4 · x · 2026-07-24
AntLingAGi has revealed the architectural details of the upcoming Ling-3.0 model, which features a native hybrid-linear attention mechanism.
The core design involves stacking KDA layers (for fine-grained control over long-range memory) with MLA layers at a 5:1 ratio. This hybrid approach not only optimizes long-context handling but also significantly boosts MoE (Mixture of Experts) compute efficiency through a 1/64 expert activation ratio.
According to the post, Ling-3.0 natively supports a 256K context length, with the architectural capability to scale up to 1M.
More from Models
- Repligate says Claude Opus 3 appears to evolve without changing its weights — repligate · 2026-07-27
- “Opus 5” post lands as a rebenchmarking-at-scale AI joke — kalomaze · 2026-07-27
- Top models now write worse than a year ago, critic says — dbreunig · 2026-07-27
- MPT-30B radar charts became an unexpectedly controversial design choice — code_star · 2026-07-27
- Local Gemma 4 31B starts acting sarcastic and users cannot reproduce it — n0head_r · 2026-07-27
- Google’s Gemini 3.6 Flash could win by matching Sonnet quality at a lower cost — haider1 · 2026-07-27