ExploreNet Learns Adaptive Exploration for Diffusion GRPO, Replacing Isotropic Noise

StellaLisy · x · 2026-10-08

The author notes that GRPO in flow-matching models adds identical isotropic Gaussian noise at every denoising step, even though latent elements differ in sensitivity — leading to exploration of unuseful directions.

ExploreNet learns an adaptive exploration distribution conditioned on the current state, rewarded by rollout diversity, enabling faster and more targeted learning in diffusion RL.

Related event: ExploreNet learns directional exploration noise to outperform FlowGRPO in diffusion RL(9 posts)→

Original post →

More from Multimodal

Multimodal channel →