Inclusion AI launches LLaDA2.2-flash, a diffusion model for agentic workloads
heyshrutimishra · x · 2026-07-25
- Inclusion AI has released LLaDA2.2-flash, a large-scale diffusion language model designed specifically for agentic workloads.
- Instead of generating tokens strictly left-to-right, it produces tokens in parallel blocks and then uses Levenshtein Editing to delete defects and insert fixes mid-sequence.
- The paper says the model is competitive with top autoregressive baselines on agent benchmarks, while running 1.6× faster.
- It also supports 128K context natively and is open source.
- The authors argue this is the first time a diffusion model has been practical for production-grade agent workloads.
Related event: inclusionAI Launches LLaDA2.2-flash: 100B Diffusion Model for Agent Latency(9 posts)→
More from Models
- Laguna S 2.1 says users want open source, simpler agentic coding stacks — max_paperclips · 2026-07-25
- Opus 5 is said to discuss honesty 6× more than other agents in Village — bronzeagepapi · 2026-07-25
- Opus 5 is catching bugs introduced by Opus 4.8 — damnGruz · 2026-07-25
- AMD open-sources Instella 16B MoE with checkpoints from pretraining to RL — bronzeagepapi · 2026-07-25
- Google posts a 1-hour agentic engineering course covering memory, MCP and multi-agent systems — ifioknkem · 2026-07-25
- Opus 5 is claimed to jump ahead on three spreadsheet-agent benchmarks — surmenok · 2026-07-25