FULL STORY

Ant's Ling-3.0-flash: From Launch to Open Source

Ant Group introduced Ling-3.0-flash, a 124B hybrid reasoning MoE model for production agents, and subsequently open-sourced it under the MIT license.

2026-07-24 ~ 2026-08-04 · 4 episodes · 18 posts

Episode 1 · Ant Group Releases Ling-3.0-flash MoE Model (2026-07-24, 12 posts)

Ant Group's inclusionAI officially released Ling-3.0-flash, a hybrid-reasoning model designed for production-grade agents. The model is now available on OpenRouter and Nous Portal, with OpenRouter offering free access until August 3, 2026, and Nous Portal providing a one-week free trial. Open weights for the model have also been confirmed.

Confirmed

Multiple sources indicate that Ling-3.0-flash is a hybrid-reasoning sparse MoE model. It has a total parameter count of 124B but only activates 5.1B parameters per token, supports a context length of up to 256K, and claims a time-to-first-token (TTFT) of less than 1 second. The model is specifically optimized for programming and agentic workflows. According to @FellMentKE, it features self-correction capabilities, allowing it to read and fix errors autonomously when generated code fails, thereby reducing developer debugging time.

Regarding deployment, Nous Portal announced a one-week free usage period, while OpenRouter's free tier extends directly to August 3, 2026. Furthermore, @victormustar confirmed that open weights for the model are planned.

Why it matters

The combination of a massive total parameter count and highly efficient parameter activation enables the model to deliver strong reasoning capabilities while significantly reducing latency and operational costs. These characteristics make it highly suitable for production-grade agent scenarios that are sensitive to response speed and expenses. The multi-year free access period and upcoming open weights provide developers with substantial convenience and ample room for testing and application.

Episode 2 · Ant Group Releases Ling-3.0-flash Hybrid Reasoning Model (2026-07-27, 2 posts)

Ant Group launched Ling-3.0-flash, a 124B parameter MoE model for production-grade agents that activates only 5.1B parameters per token, with free API access available until August 3rd.

Episode 3 · Ant Group Releases Ling-3.0-flash Hybrid Reasoning Model (2026-07-29, 2 posts)

Ant Group released Ling-3.0-flash, a native hybrid reasoning MoE model for production agents that activates only 5.1B of its 124B parameters per token to rival trillion-parameter flagships.

Episode 4 · Ling-3.0-flash Hybrid MoE Model Goes Open Source (2026-08-04, 2 posts)

The Ling-3.0-flash, a new native hybrid reasoning MoE model with over 124B parameters, has been open-sourced under the MIT license. Its FP8 version significantly lowers deployment barriers, requiring only 128GB of VRAM.