FULL STORY
Ant's Ling-3.0-flash: From Launch to Open Source
Ant Group introduced Ling-3.0-flash, a 124B hybrid reasoning MoE model for production agents, and subsequently open-sourced it under the MIT license.
2026-07-24 ~ 2026-08-04 · 4 episodes · 18 posts
Episode 1 · Ant Group Releases Ling-3.0-flash MoE Model (2026-07-24, 12 posts)
Ant Group's inclusionAI officially released Ling-3.0-flash, a hybrid-reasoning model designed for production-grade agents. The model is now available on OpenRouter and Nous Portal, with OpenRouter offering free access until August 3, 2026, and Nous Portal providing a one-week free trial. Open weights for the model have also been confirmed.
Confirmed
Multiple sources indicate that Ling-3.0-flash is a hybrid-reasoning sparse MoE model. It has a total parameter count of 124B but only activates 5.1B parameters per token, supports a context length of up to 256K, and claims a time-to-first-token (TTFT) of less than 1 second. The model is specifically optimized for programming and agentic workflows. According to @FellMentKE, it features self-correction capabilities, allowing it to read and fix errors autonomously when generated code fails, thereby reducing developer debugging time.
Regarding deployment, Nous Portal announced a one-week free usage period, while OpenRouter's free tier extends directly to August 3, 2026. Furthermore, @victormustar confirmed that open weights for the model are planned.
Why it matters
The combination of a massive total parameter count and highly efficient parameter activation enables the model to deliver strong reasoning capabilities while significantly reducing latency and operational costs. These characteristics make it highly suitable for production-grade agent scenarios that are sensitive to response speed and expenses. The multi-year free access period and upcoming open weights provide developers with substantial convenience and ample room for testing and application.
- Nous Portal opens Ling-3.0-flash free for a week, a 124B MoE model built for agents — NousResearch · 2026-07-24
- AntLing-3.0-flash launches on OpenRouter with free access through August 2026 — derspenti · 2026-07-24
- AntLing-3.0-flash goes live on OpenRouter with free access through Aug. 3, 2026 — niacolhealth · 2026-07-24
- Ant Ling launches Ling-3.0-flash, a 124B MoE model for production agents — ryanmerket · 2026-07-24
- Ling-3.0-flash launches with 124B parameters, 5.1B active, and 256K context — BanghuaZ · 2026-07-24
- Ling-3.0-flash hits OpenRouter with 256K context and confirmed open weights — victormustar · 2026-07-24
- Ling-3.0-flash pairs a 1M-token context with agent benchmarks and free access on OpenRouter — FellMentKE · 2026-07-24
- Ling-3.0-flash is pitched as a coding model that can read errors and fix its own code — FellMentKE · 2026-07-24
- Ling-3.0-flash offers 124B total parameters with just 5.1B active per token — truecakesnake · 2026-07-25
- Ant Group Releases Ling-3.0-flash: 124B Parameters with Just 5.1B Active — SonglinYang4 · 2026-07-25
- Ant’s inclusionAI launches Ling-3.0-flash with 124B parameters and 256K context — Loose_Bank1709 · 2026-07-25
- Ant Group’s inclusionAI releases Ling-3.0-flash, a 124B sparse MoE with 256K context — Loose_Bank1709 · 2026-07-25
Episode 2 · Ant Group Releases Ling-3.0-flash Hybrid Reasoning Model (2026-07-27, 2 posts)
Ant Group launched Ling-3.0-flash, a 124B parameter MoE model for production-grade agents that activates only 5.1B parameters per token, with free API access available until August 3rd.
- Ant Group launches Ling-3.0-flash with 124B parameters and a free API until Aug. 3 — Loose_Bank1709 · 2026-07-27
- Ling-3.0-flash debuts with 124B parameters and 5.1B active per token — aftahi_ai · 2026-07-27
Episode 3 · Ant Group Releases Ling-3.0-flash Hybrid Reasoning Model (2026-07-29, 2 posts)
Ant Group released Ling-3.0-flash, a native hybrid reasoning MoE model for production agents that activates only 5.1B of its 124B parameters per token to rival trillion-parameter flagships.
- Ant Group’s Ling-3.0-flash ships with 124B parameters and only 5.1B active — 智东西 · 2026-07-29
- Ling-3.0-flash Released: 5.1B Active Params Matches 1T Flagship — alifcoder · 2026-07-30
Episode 4 · Ling-3.0-flash Hybrid MoE Model Goes Open Source (2026-08-04, 2 posts)
The Ling-3.0-flash, a new native hybrid reasoning MoE model with over 124B parameters, has been open-sourced under the MIT license. Its FP8 version significantly lowers deployment barriers, requiring only 128GB of VRAM.
- Ling-3.0-flash Open-Sourced: 124B Hybrid-linear MoE Model Released — NielsRogge · 2026-08-04
- Ling-3.0-flash Open-Sourced: 127B MoE with Official FP8 Requiring Only 128GB — derspenti · 2026-08-04