Perplexity Open-Sources pplx-embed-v2-late: Late-Interaction Embedding Models for OCR-Free Multimodal Retrieval

On October 8, Perplexity released and open-sourced pplx-embed-v2-late: a pair of late-interaction multi-vector embedding models at 9B and 0.6B, built on Qwen3.5 bidirectional attention, preserving a 128-dim vector per token with MaxSim late-interaction retrieval. Both models can uniformly retrieve text, images, and full-page content—document pages such as PDFs, slides, screenshots, and charts can be retrieved directly as images without OCR. They are available on Hugging Face (MIT license per @antoinechaffin) with native Sentence Transformers support (requires sentence-transformers>=6.0.0), installable via pip.

Confirmed

Why it matters

The asymmetric retrieval paradigm built on a shared embedding space makes it feasible to combine low-cost on-device queries with large-model offline indexing; OCR-free page-level visual retrieval significantly simplifies document processing pipelines, and retrieval quality gains can propagate downstream into agent answer accuracy.

2026-10-08 ~ 2026-10-08 · 36 related posts

Primary sources

4 near-duplicate retellings: beirmug · CShorten30 · antoine_chaffin · antoine_chaffin