ExploreNet's gain comes from targeted exploration, not larger noise, ablation shows

StellaLisy · x · 2026-10-08

ExploreNet's learned noise averages 1.4x the magnitude of FlowGRPO's isotropic noise. Matching magnitude but making the noise isotropic removes the gain, proving targeted exploration — not noise size — drives the improvement. ExploreNet also works with smaller group sizes and fewer denoising steps.

Related event: ExploreNet learns directional exploration noise to outperform FlowGRPO in diffusion RL(9 posts)→

Original post →

More from Research

Research channel →