Ant Group Releases Ling-3.0-flash MoE Model

Ant Group's inclusionAI officially released Ling-3.0-flash, a hybrid-reasoning model designed for production-grade agents. The model is now available on OpenRouter and Nous Portal, with OpenRouter offering free access until August 3, 2026, and Nous Portal providing a one-week free trial. Open weights for the model have also been confirmed.

Confirmed

Multiple sources indicate that Ling-3.0-flash is a hybrid-reasoning sparse MoE model. It has a total parameter count of 124B but only activates 5.1B parameters per token, supports a context length of up to 256K, and claims a time-to-first-token (TTFT) of less than 1 second. The model is specifically optimized for programming and agentic workflows. According to @FellMentKE, it features self-correction capabilities, allowing it to read and fix errors autonomously when generated code fails, thereby reducing developer debugging time.

Regarding deployment, Nous Portal announced a one-week free usage period, while OpenRouter's free tier extends directly to August 3, 2026. Furthermore, @victormustar confirmed that open weights for the model are planned.

Why it matters

The combination of a massive total parameter count and highly efficient parameter activation enables the model to deliver strong reasoning capabilities while significantly reducing latency and operational costs. These characteristics make it highly suitable for production-grade agent scenarios that are sensitive to response speed and expenses. The multi-year free access period and upcoming open weights provide developers with substantial convenience and ample room for testing and application.

2026-07-24 ~ 2026-07-25 · 12 related posts

Full story(4 episodes)→

Primary sources

1 near-duplicate retellings: Loose_Bank1709