Why multimodal retrieval embeds rendered pages: late interaction author explains the design

antoine_chaffin · x · 2026-10-08

antoinechaffin explains two core design choices of his retrieval models:

Why multimodal: Real-world corpora are messy — PDFs, slides, scans, charts, tables. Parsing + OCR is expensive and lossy, destroying layout, tables and figures. Instead, the models directly embed the rendered page, so text queries search the page itself, and it also works on natural images.

Why late interaction: Dense models squeeze whole documents into a single vector, which breaks down as documents grow longer and more diverse (images are even denser), and a single dot product has limited expressivity. One vector per token + MaxSim lifts both bottlenecks.

Related event: Perplexity Open-Sources Multimodal Embedding Models pplx-embed-v2-late(38 posts)→

Original post →

More from Research

Research channel →