ExploreNet diverges from SD-3.5 more often (86.1%) and beats FlowGRPO on benchmarks

StellaLisy · x · 2026-10-08

Qualitative comparison against FlowGRPO: ExploreNet wins on a range of image generation benchmarks and human preference ratings, and its outputs diverge further from pretrained SD-3.5 in 86.1% of cases — attributed to better exploration. Learning dynamics show it starts worse but gains accumulate quickly through training.

Related event: ExploreNet: Learnable Exploration Distributions Improve GRPO Fine-tuning of Diffusion Models(5 posts)→

Original post →

More from Multimodal

Multimodal channel →