Latents are unevenly sensitive, so GRPO exploration should adapt per element

StellaLisy · x · 2026-10-08

GRPO in flow-matching models adds standard gaussian noise at each denoising step with an identical distribution per latent element. But latents are unevenly sensitive, so the model can end up exploring 'un-useful' directions — exploration should adapt to each element's sensitivity.

Related event: ExploreNet learns directional exploration noise to outperform FlowGRPO in diffusion RL(9 posts)→

Original post →

More from Research

Research channel →