Ling-3.0-flash Open-Sourced: 124B Hybrid-linear MoE Model Released
NielsRogge · x · 2026-08-04
AntLingAGI has open-sourced Ling-3.0-flash, a next-generation native hybrid reasoning model, on Hugging Face.
- Architecture: It utilizes a 124B-A5.1B MoE design. It leverages a Native Hybrid-Linear Architecture from the start of pre-training, incorporating Kimi Delta Attention (KDA) and Multi-Latent Attention (MLA).
- Performance: The model matches or outperforms its predecessor across key benchmarks.
- Engineering: It natively integrates the SGLang HiCache + Mooncake hierarchical caching architecture for efficient KV-cache storage. Despite its smaller activated size, it delivers impressive reasoning, instruction following, and long-context capabilities tailored for complex agentic workflows in production.
Related event: Ling-3.0-flash Hybrid MoE Model Goes Open Source(2 posts)→
More from Infra
- Valar Atomics Vision: Cheap Nuclear Energy to Power AI Robotics in Heavy Industry — johncoogan · 2026-08-05
- LMSYS Releases SpecForge Update: Advanced Speculative Decoding for Major Models — BanghuaZ · 2026-08-05
- Proposed US Ban on Chinese Optics to Drive Up AI Infrastructure Costs — tengyanAI · 2026-08-05
- Testing MiniMax H3 on RTX 5080: 15-Second Video in 11 Minutes — FreeTheClanks · 2026-08-05
- KERNEL open-sources Hypeman, sandbox infra for agentic workloads — ycombinator · 2026-08-05
- Scaling Real-Time AI Agents: Introducing Session-Aware Load Balancing — rseroter · 2026-08-05