Why multimodal retrieval embeds rendered pages: late interaction author explains the design
antoine_chaffin · x · 2026-10-08
antoinechaffin explains two core design choices of his retrieval models:
Why multimodal: Real-world corpora are messy — PDFs, slides, scans, charts, tables. Parsing + OCR is expensive and lossy, destroying layout, tables and figures. Instead, the models directly embed the rendered page, so text queries search the page itself, and it also works on natural images.
Why late interaction: Dense models squeeze whole documents into a single vector, which breaks down as documents grow longer and more diverse (images are even denser), and a single dot product has limited expressivity. One vector per token + MaxSim lifts both bottlenecks.
Related event: Perplexity Open-Sources Multimodal Embedding Models pplx-embed-v2-late(38 posts)→
More from Research
- KERNAUT uses coding agents and QD search to auto-discover interpretable kernel models — sirbayes · 2026-10-08
- Kernaut: coding agents design Gaussian process kernels via program search — sirbayes · 2026-10-08
- Kernaut's discovered kernel beats tuned standard kernels on glucose prediction — sirbayes · 2026-10-08
- Kernaut's discovered kernels stay interpretable: 16 scalar functions, inner-product form — sirbayes · 2026-10-08
- iOSWorld brings computer-use agent benchmarking to iOS at COLM 2026 — kohjingyu · 2026-10-08
- Epoch's InnovationEval: AI agents still far from producing real research innovations — Afinetheorem · 2026-10-08