ExploreNet outperforms FlowGRPO by learning state-dependent exploration noise
StellaLisy · x · 2026-10-08
StellaLisy introduces ExploreNet: for GRPO fine-tuning of flow-matching models, it learns a state-dependent exploration noise distribution to maximize reward spread instead of fixed isotropic gaussian noise. It beats the FlowGRPO baseline across image generation benchmarks and human preference ratings, with gains accumulating quickly over training.
Related event: ExploreNet Learns State-Dependent Exploration Noise, Beating FlowGRPO(2 posts)→
More from Multimodal
- GPT-6 Astra nails 3D: reads user Memory to build poster from Blender models — op7418 · 2026-10-09
- Prompt Templates for AI Video: Background and Object Replacement That Actually Hold — ifioknkem · 2026-10-09
- Prompt Templates for AI Video Editing: Background and Object Replacement — ifioknkem · 2026-10-09
- Gemini Prompt Templates: Outfit Swaps and Camera Angle Changes That Keep Subjects Consistent — ifioknkem · 2026-10-09
- Gemini Can Now Edit Your Videos: 10 Prompts for Outfit Swaps, Object Removal and More — ifioknkem · 2026-10-09
- ExploreNet ablation: perturbing learned high-sensitivity channels drives larger human-perceived change — StellaLisy · 2026-10-08