Perplexity Open-Sources pplx-embed-v2-late: Late-Interaction Embedding Models for OCR-Free Multimodal Retrieval
On October 8, Perplexity released and open-sourced pplx-embed-v2-late: a pair of late-interaction multi-vector embedding models at 9B and 0.6B, built on Qwen3.5 bidirectional attention, preserving a 128-dim vector per token with MaxSim late-interaction retrieval. Both models can uniformly retrieve text, images, and full-page content—document pages such as PDFs, slides, screenshots, and charts can be retrieved directly as images without OCR. They are available on Hugging Face (MIT license per @antoinechaffin) with native Sentence Transformers support (requires sentence-transformers>=6.0.0), installable via pip.
Confirmed
- The 9B and 0.6B share the same embedding space: you can encode an entire corpus offline with the 9B and serve online queries with the 0.6B. Both models were distilled from a single 18B teacher via aligned per-token embeddings.
- Multimodal results: in ViDoRe v3 image mode, the 9B averages 65.2% (nDCG@10), beating nemotron-colembed-v2-8b and trailing only the vision-specialized EVIE, despite a much smaller embedding dimension; the 0.6B scores 62.3%, just 1.2 points behind the 8B Nemotron ColEmbed V2. Per @antoinechaffin, the 9B outperforms rivals on MIRACL-Vision except Gemini-Emb (vision-specialized), and the 0.6B comes close to Qwen3-VL-Embed-8B.
- Agentic retrieval: on BrowseComp+, a GPT-OSS-120B agent paired with the 9B retriever reaches 64.0% accuracy, 4.9 points above the next-best ColBERT model, while making fewer search calls; it was also evaluated on the MADQA multimodal task.
- Cross-model retrieval tested: querying with the 0.6B against a 9B-indexed document collection outperforms 0.6B on both sides, closing roughly half the gap to the 9B-9B combo with no added query latency—offering a production accuracy-latency trade-off.
- 0.6B details (@tomaarsen): starting from Qwen3.5-0.8B, the text tower was pruned from 24 to 12 layers; total params are 594M, with roughly 240M active for text encoding and 340M for images.
Why it matters
The asymmetric retrieval paradigm built on a shared embedding space makes it feasible to combine low-cost on-device queries with large-model offline indexing; OCR-free page-level visual retrieval significantly simplifies document processing pipelines, and retrieval quality gains can propagate downstream into agent answer accuracy.
2026-10-08 ~ 2026-10-08 · 36 related posts
Primary sources
- Perplexity releases pplx-embed-v2-late: OCR-free late-interaction embeddings topping retrieval benchmarks — perplexity_ai ·
- Perplexity's pplx-embed-v2-late: serve 0.6B queries against 9B document embeddings — tomaarsen ·
- On multimodal retrieval, the 9B scores 65.2% on ViDoRe v3, beating nemotron-colembed-v2-8b — antoine_chaffin ·
- [source] Perplexity releases pplx-embed-v2-late: OCR-free late-interaction embeddings topping retrieval benchmarks — perplexity_ai · 2026-10-08
- Perplexity ships pplx-embed-v2 embeddings, tops MADQA retrieval at 92.4% — perplexity_ai · 2026-10-08
- Open 9B ColBERT retrieval models released, frontier results on text, multimodal and agentic retrieval — antoine_chaffin · 2026-10-08
- Contextual AI releases 9B and 0.6B late-interaction embedding models hitting frontier retrieval results — antoine_chaffin · 2026-10-08
- Why multimodal retrieval embeds rendered pages: late interaction author explains the design — antoine_chaffin · 2026-10-08
- 18B teacher distilled into 9B/0.6B students sharing one multi-vector embedding space — antoine_chaffin · 2026-10-08
- Asymmetric retrieval: index with 9B, encode queries with 0.6B for frontier accuracy at low latency — antoine_chaffin · 2026-10-08
- Query a 9B index with the 0.6B model: half the accuracy gap at zero query-time cost — antoine_chaffin · 2026-10-08
- Trained on 186M pairs from 594 datasets in 46 languages, with benchmark data fully removed — antoine_chaffin · 2026-10-08
- Top 1 and 2 on Q2D-Web: the 0.6B dethrones former leader nemotron-embed-8b — antoine_chaffin · 2026-10-08
- [source] On multimodal retrieval, the 9B scores 65.2% on ViDoRe v3, beating nemotron-colembed-v2-8b — antoine_chaffin · 2026-10-08
- 9B tops 72 MTEB-style tasks; 0.6B matches gemini-embedding-2 on text retrieval — antoine_chaffin · 2026-10-08
- 0.6B visual embedding model rivals 8B models, keeps natural image search strong — antoine_chaffin · 2026-10-08
- Cross-model retrieval tested: 0.6B querying a 9B index closes half the gap at no latency cost — antoine_chaffin · 2026-10-08
- Cross-model retrieval: pplx-embed-v2's 0.6B querying a 9B index beats a 0.6B index — antoine_chaffin · 2026-10-08
- pplx-embed-v2 agentic retrieval evals: 9B model hits 64.0% on BrowseComp-Plus — antoine_chaffin · 2026-10-08
- Perplexity open-sources pplx-embed-v2 embeddings, with 9B model hitting 92.4% on multimodal retrieval — antoine_chaffin · 2026-10-08
- More details on the open retrieval models: asymmetric design for production, SOTA text retrieval too — antoine_chaffin · 2026-10-08
- Perplexity open-sources pplx-embed-v2-late multimodal embedding models — CShorten30 · 2026-10-08
- Perplexity releases pplx-embed-v2-late: 0.6B & 9B text+image retrievers — tomaarsen · 2026-10-08
- Perplexity's pplx-embed-v2-late retrieves PDFs and screenshots as images, no OCR needed — tomaarsen · 2026-10-08
- pplx-embed-v2-late built on Qwen3.5, keeps one 128d vector per token with MaxSim — tomaarsen · 2026-10-08
- [source] Perplexity's pplx-embed-v2-late: serve 0.6B queries against 9B document embeddings — tomaarsen · 2026-10-08
- 9B and 0.6B Embedding Models Share One Vector Space for Encode-Small-Search — tomaarsen · 2026-10-08
- Shared embedding space: index with 9B, serve queries with 0.6B — tomaarsen · 2026-10-08
- 0.6B trails an 8B model by just 1.2 points on ViDoRe v3 page images — tomaarsen · 2026-10-08
- pplx-embed 9B lifts agent accuracy to 64%, +4.9 over next best on BrowseComp+ — tomaarsen · 2026-10-08
- 0.6B model starts from Qwen3.5-0.8B, pruning text tower to 12 layers — tomaarsen · 2026-10-08
- Perplexity's 0.6B embedding model scores 62.3 on ViDoRe v3, beating larger rivals — antoine_chaffin · 2026-10-08
- Perplexity open-sources pplx-embed-v2 embedding models under MIT license — tomaarsen · 2026-10-08
- Perplexity's MIT-licensed ColPali models now natively supported in sentence-transformers — antoine_chaffin · 2026-10-08
- Perplexity's Embedding Models Share a Vector Space, Enabling Cross-Model Retrieval — antoine_chaffin · 2026-10-08
4 near-duplicate retellings: beirmug · CShorten30 · antoine_chaffin · antoine_chaffin