ExploreNet Learns Adaptive Exploration for Diffusion GRPO, Replacing Isotropic Noise
StellaLisy · x · 2026-10-08
The author notes that GRPO in flow-matching models adds identical isotropic Gaussian noise at every denoising step, even though latent elements differ in sensitivity — leading to exploration of unuseful directions.
ExploreNet learns an adaptive exploration distribution conditioned on the current state, rewarded by rollout diversity, enabling faster and more targeted learning in diffusion RL.
More from Multimodal
- ComfyUI app ANIMA goes viral: Qwen 3.5 'hallucinates' popular photos, 300K views — Ok_Contribution8157 · 2026-10-09
- NAMVIS (NeurIPS 2026): next-scale autoregression beats diffusion for novel-view synthesis, 3x faster — RexDouglass · 2026-10-09
- fal Engineer on Video Speed: H3 Max Renders 15 Seconds of Video in 5 — OdinLovis · 2026-10-09
- Odyssey launches Odyssey-3, claims SOTA world model on Physics-IQ benchmark — Scobleizer · 2026-10-09
- One prompt, one full movie: dev shows multi-director agent pipeline for AI films — Exciting-Income-5840 · 2026-10-09
- Nano Banana 2.1 becomes Google's best image editing model across all 7 edit actions — ArtificialAnlys · 2026-10-09