Ling-3.0 Architecture Revealed: Native Hybrid Attention with 256K Context

SonglinYang4 · x · 2026-07-24

AntLingAGi has revealed the architectural details of the upcoming Ling-3.0 model, which features a native hybrid-linear attention mechanism.

The core design involves stacking KDA layers (for fine-grained control over long-range memory) with MLA layers at a 5:1 ratio. This hybrid approach not only optimizes long-context handling but also significantly boosts MoE (Mixture of Experts) compute efficiency through a 1/64 expert activation ratio.

According to the post, Ling-3.0 natively supports a 256K context length, with the architectural capability to scale up to 1M.

Original post →

More from Models

Models channel →