ExploreNet learns directional exploration noise to outperform FlowGRPO in diffusion RL

TL;DR: StellaLi, Xiaochuang Han, Yulia Tsvetkov, Luke Zettlemoyer, and others have released the paper "ExploreNet: Learning Where to Explore in Diffusion GRPO," which turns the exploration noise in diffusion model GRPO fine-tuning from a fixed isotropic Gaussian into a learnable, state-adaptive distribution, letting RL training itself learn "where to explore." Author @StellaLisy posted multiple threads explaining the work, stating that the method beats FlowGRPO on multiple image generation benchmarks and human preference scores. A key takeaway: exploration gains come from direction, not magnitude.

Confirmed

Why it matters

2026-10-08 ~ 2026-10-08 · 9 related posts

Primary sources

1 near-duplicate retellings: StellaLisy