Flash-BoN: A Stronger Baseline for Diffusion Models
RisingSayak · x · 2026-07-14
The authors argue that for test-time scaling in diffusion models, best-of-N (BoN)—when done right—is a stronger baseline than many fancier methods.
Key points:
- Looking solely at NFE creates an illusion of efficiency because it ignores verifier costs.
- They propose Flash-BoN, encouraging broader sampling exploration over repeated intermediate verification.
- Evaluations should report both NFE and wall-clock time; otherwise, algorithms with frequent verification might be falsely deemed more efficient.
- Conclusion: Spending test-time compute on exploring more candidates allows BoN to often match or outperform more complex methods.
The original post includes an extended thread, noting that this is an empirical conclusion regarding how test-time compute should be evaluated.
Related event: Flash-BoN: A Stronger Baseline for Diffusion Inference and Post-Training(7 posts)→
More from Research
- Why a 1GW Chinese AI data center may be plausible after all — teortaxesTex · 2026-07-22
- Chinese AI labs are now treating distillation obfuscation as the top research topic — pmddomingos · 2026-07-22
- Structural ensembles beat single predictions in TCR:pMHC generalization study — quaidmorris · 2026-07-22
- RSS launches under OMSF to push structural biology data modeling at scale — MoAlQuraishi · 2026-07-22
- enFoldX tops 8 neoantigen scans and an unseen-peptide benchmark — quaidmorris · 2026-07-22
- enFoldX reaches AUC 0.82 on human VDJdb and transfers to mouse at 0.76 — quaidmorris · 2026-07-22