ExploreNet Boosts Diffusion GRPO by 14% by Learning Where to Explore
StellaLisy · x · 2026-10-08
The arXiv paper ExploreNet: Learning Where to Explore in Diffusion GRPO (Li, Han, Tsvetkov, Zettlemoyer) shows isotropic Gaussian noise in Flow-GRPO wastes exploration: latent elements differ wildly in how much they change the image.
- ExploreNet predicts a per-element noise scale from the latent, denoising step, and prompt before any reward is observed; it trains on group reward spread and is discarded after training, leaving inference unchanged
- On Stable Diffusion 3.5 Medium it improves held-out GenEval2 by 14% over Flow-GRPO, transfers to two compositional benchmarks and five preference/quality models, and hits 67.2% human preference win-rate
- Key findings: exploration is learnable, the shape of the exploration distribution outweighs its magnitude, and rollout quality beats quantity
- A human study confirms annotators perceive larger visual changes when channels ExploreNet amplifies are perturbed
More from Multimodal
- GPT-6 Astra nails 3D: reads user Memory to build poster from Blender models — op7418 · 2026-10-09
- Prompt Templates for AI Video: Background and Object Replacement That Actually Hold — ifioknkem · 2026-10-09
- Prompt Templates for AI Video Editing: Background and Object Replacement — ifioknkem · 2026-10-09
- Gemini Prompt Templates: Outfit Swaps and Camera Angle Changes That Keep Subjects Consistent — ifioknkem · 2026-10-09
- Gemini Can Now Edit Your Videos: 10 Prompts for Outfit Swaps, Object Removal and More — ifioknkem · 2026-10-09
- ExploreNet ablation: perturbing learned high-sensitivity channels drives larger human-perceived change — StellaLisy · 2026-10-08