Decoding alone moves diffusion retrieval Hit@1 by 6.6-13.7 points, not the paradigm
_reachsumit · x · 2026-10-07
This arXiv paper challenges the simple narrative that diffusion beats or loses to autoregressive models in generative retrieval: recent work swapping autoregressive decoders for diffusion also changed identifiers, training recipes, and decoding at once, so reported gaps cannot be attributed to the paradigm.
Fixing identifier length and training budget, the authors train autoregressive, masked-diffusion, and block-diffusion retrievers on NQ320K and MS300K with residual-quantised, product-quantised, and random identifiers, then decode each model in multiple ways. Key findings:
- Decoding alone shifts a diffusion model's Hit@1 by 6.6–13.7 points, meaning much of the reported gap comes from decoding, not paradigm.
- A reference decoding scheme, generate-and-match, generates an identifier then retrieves the nearest corpus identifiers; the generated identifier is correct for 14–21% of NQ320K queries.
- One-pass scoring — reading a fully masked identifier once and scoring each document by its codes' probabilities — matches or beats generate-and-match in 11 of 12 settings, and removes 46–83% of masked diffusion's deficit to beam search when starting from one sampled identifier.
- Autoregressive models still lead in Hit@1; on NQ320K the lead comes from the model, not beam search.
- On NQ320K, every paradigm largely memorizes which identifier answers which query: random identifiers retain 83–90% of the Hit@1 of residual-quantised ones.
More from Research
- Andrew Davison: robots need object-based SLAM, not scan-then-fit reconstructions — AjdDavison · 2026-10-07
- CtrlCache Speeds Up Interactive Video World Models 1.21–1.41x Without Retraining — Shangye Song · 2026-10-07
- Training-Free Accent Analogy Guidance Boosts Speaker Similarity in Cross-Lingual Voice Cloning — Yoomee Cho · 2026-10-07
- Source Attribution of Synthetic Data Hits 98.7% Accuracy but Falls to 29% After Style Rewriting — Joss Armstrong · 2026-10-07
- Physicist finds fractal patterns (D 1.3-1.5) cut stress response by up to 60% — aakashgupta · 2026-10-07
- AI has now cracked at least 10 open math problems each worthy of a Fields Medal — luismbat · 2026-10-07