FIRM needs only 10 NFEs and 0.24s per image on CelebA deblurring, beating baselines

prof_kamilov · x · 2026-10-07

FIRM turns the theory into an end-to-end trained flow: with N=2 sampling steps it uses just 10 network evaluations and no inference-time guidance. On CelebA deblurring it reconstructs an image in 0.24s (vs. 13.03s for baselines, 50x fewer NFEs) with higher PSNR, recovering detail baselines miss in only 2 ODE steps. A practical bonus: varying the number of sampling steps controls the distortion-perception trade-off—fewer steps favor reconstruction accuracy, more steps recover richer perceptual detail—without retraining, so one model spans the whole range. Authors: Shirin Shoushtari, Edward Chandler, Xiao Shi, Ulugbek Kamilov (WashU / UW-Madison); paper and code are available.

Related event: FIRM Cuts Network Evaluations 50x for Inverse Imaging(3 posts)→

Original post →

More from Research

Research channel →