Flash-BoN: cheap drafts beat guided search for diffusion inference-time scaling, +8% AUC at scale
gowthami_s · x · 2026-09-10
Gowthami Somepalli's team presents Flash-BoN at ECCV, showing that wall-clock cost—not NFE count—is the right yardstick for diffusion inference-time scaling.
- Under matched end-to-end runtime, simple Best-of-N already matches or beats several guided search methods, since repeated intermediate verification eats the time budget that could explore new candidates.
- Flash-BoN jointly optimizes three acceleration knobs—timestep truncation, layer skipping and activation proxies—once per model to mass-produce cheap drafts, then prunes with pointwise scores, head-to-head compares survivors, and refines the winner at full quality.
- Across 9 model-benchmark combinations, Flash-BoN leads normalized AUC everywhere, with gains growing to +8% AUC at larger model scales.
- The idea transfers to RL post-training: draft 16 rollouts, train on the top 6 and bottom 2.
More from Multimodal
- 3D Gaussian splat reconstruction of Amazon Prime Air crash site from NTSB footage — bilawalsidhu · 2026-09-10
- 3D Gaussian Splat Rebuilds Amazon Prime Air Crash Site from New NTSB Footage — bilawalsidhu · 2026-09-10
- OpenAI case study: GPT-6 Astra builds a house in Blender from one prompt, then moves it into UE5 — xiaohu · 2026-09-10
- OpenAI demos Astra driving Blender to build editable 3D scenes from prompts — xiaohu · 2026-09-10
- One prompt to an interactive UE5 home: OpenAI's GPT-6 Astra 3D design case — xiaohu · 2026-09-10
- VivagoR1 launches as a conversational AI video agent promising deterministic 5-minute outputs — 量子位 · 2026-09-10