ExploreNet outperforms FlowGRPO by learning state-dependent exploration noise

StellaLisy · x · 2026-10-08

StellaLisy introduces ExploreNet: for GRPO fine-tuning of flow-matching models, it learns a state-dependent exploration noise distribution to maximize reward spread instead of fixed isotropic gaussian noise. It beats the FlowGRPO baseline across image generation benchmarks and human preference ratings, with gains accumulating quickly over training.

Related event: ExploreNet Learns State-Dependent Exploration Noise, Beating FlowGRPO(2 posts)→

Original post →

More from Multimodal

Multimodal channel →