ExploreNet: learning targeted exploration noise for GRPO in flow-matching models

StellaLisy · x · 2026-10-08

The authors explain ExploreNet's motivation: GRPO in flow-matching models adds identical gaussian noise per denoising step, but latents are unevenly sensitive, leading to 'useless' exploration directions. Instead of cherry-picking useful rollouts (which wastes inference compute), they formulate exploration learning as an RL problem — predicting the noise distribution from the current state to maximize reward spread.

Related event: ExploreNet learns directional exploration noise to outperform FlowGRPO in diffusion RL(9 posts)→

Original post →

More from Research

Research channel →