ExploreNet learns directional exploration noise to outperform FlowGRPO in diffusion RL
TL;DR: StellaLi, Xiaochuang Han, Yulia Tsvetkov, Luke Zettlemoyer, and others have released the paper "ExploreNet: Learning Where to Explore in Diffusion GRPO," which turns the exploration noise in diffusion model GRPO fine-tuning from a fixed isotropic Gaussian into a learnable, state-adaptive distribution, letting RL training itself learn "where to explore." Author @StellaLisy posted multiple threads explaining the work, stating that the method beats FlowGRPO on multiple image generation benchmarks and human preference scores. A key takeaway: exploration gains come from direction, not magnitude.
Confirmed
- Motivation: running GRPO on flow-matching models typically injects standard Gaussian noise at every denoising step, with the same noise distribution for all latent elements—but element sensitivities are uneven, so the model may explore "useless" directions.
- Method: ExploreNet predicts an adaptive exploration distribution based on the current state, trained with reward dispersion as the objective; unlike selecting more useful rollouts within a group, it wastes no extra inference compute.
- Results: with FlowGRPO (isotropic exploration noise) as the baseline, ExploreNet performs better across multiple image generation benchmarks and human preference scores, and its generated images deviate further from the baseline 86.1% of the time.
- Ablation 1 (magnitude alignment): the noise learned by ExploreNet has 1.4x the average magnitude of isotropic Gaussian noise, but aligning the magnitude and reverting to isotropic noise eliminates the gains—proving the key is "directed exploration," not louder noise.
- Ablation 2 (channel sensitivity): the model assigns larger noise scales to latent channels that "change the image more," so exploration automatically adapts to the sensitivity structure of the latent space.
- Ablation 3 (human perception): when annotators rated images with a single latent channel perturbed, the channels the model learned to amplify produced noticeably larger perceptual differences, showing the learned exploration directions truly align with the goal of "changing the image."
Why it matters
- This work makes "exploration" itself the object of RL learning rather than relying on hand-designed noise priors, offering a new paradigm for RL fine-tuning of diffusion models.
- The finding that "shape matters more than magnitude" offers direct guidance for designing exploration mechanisms in future diffusion policies.
2026-10-08 ~ 2026-10-08 · 9 related posts
Primary sources
- ExploreNet Learns Adaptive Exploration for Diffusion GRPO, Replacing Isotropic Noise — StellaLisy · 2026-10-08
- Latents are unevenly sensitive, so GRPO exploration should adapt per element — StellaLisy · 2026-10-08
- [source] ExploreNet outperforms FlowGRPO by learning state-dependent exploration noise — StellaLisy · 2026-10-08
- ExploreNet diverges from SD-3.5 more often (86.1%) and beats FlowGRPO on benchmarks — StellaLisy · 2026-10-08
- ExploreNet's gain comes from targeted exploration, not larger noise, ablation shows — StellaLisy · 2026-10-08
- ExploreNet assigns larger noise scales to channels that change the image more — StellaLisy · 2026-10-08
- ExploreNet ablation: perturbing learned high-sensitivity channels drives larger human-perceived change — StellaLisy · 2026-10-08
- [source] ExploreNet Boosts Diffusion GRPO by 14% by Learning Where to Explore — StellaLisy · 2026-10-08
1 near-duplicate retellings: StellaLisy