FIRM needs only 10 NFEs and 0.24s per image on CelebA deblurring, beating baselines
prof_kamilov · x · 2026-10-07
FIRM turns the theory into an end-to-end trained flow: with N=2 sampling steps it uses just 10 network evaluations and no inference-time guidance. On CelebA deblurring it reconstructs an image in 0.24s (vs. 13.03s for baselines, 50x fewer NFEs) with higher PSNR, recovering detail baselines miss in only 2 ODE steps. A practical bonus: varying the number of sampling steps controls the distortion-perception trade-off—fewer steps favor reconstruction accuracy, more steps recover richer perceptual detail—without retraining, so one model spans the whole range. Authors: Shirin Shoushtari, Edward Chandler, Xiao Shi, Ulugbek Kamilov (WashU / UW-Madison); paper and code are available.
Related event: FIRM Cuts Network Evaluations 50x for Inverse Imaging(3 posts)→
More from Research
- QUEEN paper distills AlphaZero-style chess search into language for LLMs — danqi_chen · 2026-10-07
- COLM paper asks whether VLMs can internalize tool calls in latent space instead of calling them — PMinervini · 2026-10-07
- CLeaR 2027 Opens Its Submission Portal — ArthurGretton · 2026-10-07
- MIT paper: a minimalist agent loop that passes history as code variables beats Letta and ACE at half the cost — rohanpaul_ai · 2026-10-07
- doodlestein releases explainer video alongside open-source companion repo — doodlestein · 2026-10-07
- A 3Blue1Brown-Style Explainer Rendered End-to-End in Rust by franken_manim — doodlestein · 2026-10-07