Flash-BoN Also Accelerates Post-Training

RisingSayak · x · 2026-07-14

The author adds that the "cheap candidate generation" in Flash-BoN is useful not only for test-time scaling but also for accelerating RL post-training.

They share their results: Flash-Flow-GRPO achieves the convergence of the baseline Flow-GRPO in 10× fewer steps. In the replies, the author also notes that Flash-BoN can be combined with other methods and is entirely complementary.

Related event: Flash-BoN: A Stronger Baseline for Diffusion Inference and Post-Training(7 posts)→

Original post →

More from Research

Research channel →