PreviewDiff: Multimodal Critic-Guided Search over Diffusion Latents
Vighnesh Subramaniam, Boris Katz, Brian Cheung, Chun-Liang Li, Tomas Pfister, Yale Song
cs.CV, cs.AI, cs.LG
2026-09-29
A multimodal critic critiques mid-denoise previews; its fixes branch the trajectory, and this training-free search beats budget-matched Best-of-N on image and video benchmarks.
Diffusion models render convincingly but keep failing on compositional detail: wrong object counts, attributes bound to the wrong subject, reversed spatial relations, actions drifting off the prompt's timeline. Aggregate metrics like FID or CLIPScore rate such outputs as well aligned; targeted suites such as GenEval catch the failures. The standard fix is more test-time compute. Best-of-N sampling draws N complete outputs and lets a reward model or multimodal judge pick the best one. It is cheap and ubiquitous, and it is open-loop: judgment happens after generation, a promising trajectory that has already drifted cannot be repaired, and compute spent on samples that fail late is written off. The paper asks whether the same verifier budget buys more when spent inside denoising.
PreviewDiff is training-free and wraps a frozen generator. Intermediate denoising states become nodes of a search tree.
The design choices are argued, not decorative. Corrections are appended as notes rather than full prompt rewrites, because a rewrite tends to drop constraints the partial sample already satisfies. Re-noising is local, a few steps back rather than back to pure noise, so a branch stays near the current scene while leaving room for the fix to land. The tree policy is a beam instead of UCT visit counts, since every expansion costs a decode plus one judge call and nothing can be visited twice cheaply. In MCTS terms: states are previews, actions are semantic corrections, rewards are prompt-satisfaction scores, backup is a max over children.
Scaling first, mean Gemini score on a 0–4 rubric. On SDXL over GenEval2, PreviewDiff rises from 2.17 to 2.86 as the budget grows from 86 to 704 denoiser steps per prompt; budget-matched Best-of-N manages 1.91 to 2.71 and trails at every budget. On LTX-Video over NarrLV the gap at the largest budget is +0.26 (2.39 versus 2.13).
Fixed-budget comparisons across backbones:
| Setting | Method | Metric | Result |
| SDXL / GenEval2 | PreviewDiff vs Best-of-N | Gemini P@4 | 2.292 vs 2.090 |
| SDXL / GenEval2 | EvoSearch / Particle | Gemini P@4 | 2.103 / 1.902 |
| FLUX.2 9B / GenEval2 | PreviewDiff vs Best-of-N | Gemini P@4 | 3.112 vs 2.775 |
| LTX-Video / NarrLV | PreviewDiff vs Best-of-N | Gemini P@4 | 2.167 vs 1.544 |
| LTX-Video / T2V-CompBench | PreviewDiff vs Video-T1 | Gemini P@4 | 3.021 vs 2.861 |
| Wan 2.2 / VBench 2.0 | PreviewDiff vs Best-of-N | Gemini P@4 | 3.554 vs 3.292 |
The gains survive a change of evaluator: Qwen-3, BLIP-VQA, LLaVA and Qwen-2.5 columns all move up, and running the search with Qwen-3 as the in-loop critic also improves results, so this is not an artifact of the Gemini rubric. Nine human raters judged 360 image and 200 video prompts; PreviewDiff took the largest three-way preference share and won both pairwise matchups (against EvoSearch and Best-of-N) above chance, with content gains clearly larger than style gains.
Ablations: search width is the biggest lever, taking SDXL from 1.97 to 2.74 and LTX-Video from 1.64 to 2.30 as width grows from 1 to 8. Tree depth (1–3) and semantic variants (1–4) add smaller amounts. Earlier checkpoints beat later ones; previews only need to be legible enough for the critic to spot a plausible fix. Removing language feedback and keeping bare scores drops SDXL to 2.063 and LTX-Video to 2.014.
Test-time scaling is routine for language models, while visual generation has mostly settled for Best-of-N. This paper relocates the verifier from judge at the end to controller inside the process, and shows the same budget works harder there, holding from 2.5B to 9B backbones. That makes it complementary to model scaling rather than redundant. Practically, it is training-free and fits any diffusion model that exposes intermediate clean-latent estimates, so quality-critical, latency-tolerant applications can adopt it directly. One caveat belongs here rather than in the results: budgets are matched in denoiser steps and judge calls, not wall-clock. Best-of-N samples run in parallel; a tree has sequential dependencies, and real latency will diverge by more than the step counts suggest.
The authors' own list: the critic optimizes prompt adherence only, so style gains are weak (human evaluation confirms it), and the action space has to stay finite because each expansion costs a decode plus a judge call. Three further issues. Wall-clock latency is never reported. The in-search critic and the primary metric share one Gemini model and rubric, and GenEval2 stresses exactly the error types the critic is told to hunt for; auxiliary evaluators blunt the concern without removing it. And the prompt-editing ablation is nearly flat, 2.292 dropping only to 2.285 without edits, which sits oddly beside the claim that semantic corrections are the core mechanism. Most of the lift comes from width and branching. The two ablations that strip language feedback (2.063 versus 2.285) are never reconciled, and a restart-depth ablation promised in the methods section never appears.