ExploreNet: learning targeted exploration noise for GRPO in flow-matching models
StellaLisy · x · 2026-10-08
The authors explain ExploreNet's motivation: GRPO in flow-matching models adds identical gaussian noise per denoising step, but latents are unevenly sensitive, leading to 'useless' exploration directions. Instead of cherry-picking useful rollouts (which wastes inference compute), they formulate exploration learning as an RL problem — predicting the noise distribution from the current state to maximize reward spread.
More from Research
- Tencent's STEPQuant: 6-bit recurrent states match FP32 with 68.7% less memory — _akhaliq · 2026-10-09
- Stanford's Level-of-Token Diffusion cuts image and video generation cost with multiresolution tokens — GordonWetzstein · 2026-10-09
- Datapoint launches Streams: live human feedback RL for multimodal models — _akhaliq · 2026-10-09
- Agent Plasticity: top-performing agents aren't the most efficient learners — RulinShao · 2026-10-09
- BiGym 2.0: LLM-Written Policies Hit 65% From One Demo, But π0.5 Still Leads at 75% — stepjamUK · 2026-10-09
- Jonathon Stray seeks engineer for large-scale study of long-term AI use on young adults — RishiBommasani · 2026-10-09