Sakana says multiple diffusion models plus MCTS beat test-time scaling on coding and math
SakanaAILabs · x · 2026-07-21
UnMaskFork uses multiple diffusion language models and MCTS to improve coding and math
Sakana AI says its ICML 2026 paper, UnMaskFork, explores whether test-time scaling can work for masked diffusion language models (MDLMs).
- Instead of increasing randomness with temperature, the method creates diversity through model switching.
- Multiple MDLMs collaborate to unmask a single answer, while Monte Carlo Tree Search searches for promising generation paths.
- The approach requires no extra training and no model changes; it works at inference time by combining pre-trained models.
- The team says the method consistently outperforms existing test-time scaling baselines on coding benchmarks and also scales well on math tasks.
- Sakana frames the work as part of its broader “collective intelligence of AI” line of research, alongside AB-MCTS and Sakana Fugu.
Related event: Sakana AI Enhances Reasoning with Collaborative Diffusion Models(2 posts)→
More from Models
- BullshitBench update: GPT-6-Astra beats all prior OpenAI models but still trails Anthropic — scaling01 · 2026-09-11
- Astra Scores 83% on GauntletBench, First Computer-Use Agent to Beat Human Baseline — ducha_aiki · 2026-09-11
- Kimi K2.8 Preview rolls out: near-K3 coding performance, 1M context for all tiers — teortaxesTex · 2026-09-11
- Looking for a classifier of software engineering task shapes to pick models per task — StewartalsopIII · 2026-09-11
- DeepSeek V4 Pro API to continue after Sept 2026, billing unchanged — teortaxesTex · 2026-09-11
- DeepSeek V4.1 Flash Hits 98% of GPT-6 Astra's Score at 1.4% of the Cost in Third-Party Benchmark — ayushtweetshere · 2026-09-11