Perplexity releases pplx-embed-v2-late: OCR-free late-interaction embeddings topping retrieval benchmarks
perplexity_ai · x · 2026-10-08
Perplexity has open-sourced pplx-embed-v2-late on Hugging Face: two late-interaction embedding models that retrieve text, images, and rendered pages in a shared embedding space.
How it works
- Unlike dense embeddings that compress a document into one vector, each token keeps a 128-dim vector scored with MaxSim, so query tokens match their closest document tokens.
- Images and rendered pages are embedded directly, so PDFs, slides, and scans are searchable without OCR, preserving tables, figures, and layout.
Benchmarks
- 92.4% on MADQA (500 questions over 800 PDFs), the top retriever score.
- On Q2D-Web (70k production queries over 190M pages), the 9B and 0.6B models hit 74.8% and 73.6% Recall@1000 vs 69.3% for nemotron-embed-8b.
- Agentic search: a GPT-OSS-120B agent with the 9B model answers 64.0% of BrowseComp+ questions, 4.9 points ahead of ColBERT, with fewer searches than any baseline.
- Both sizes are distilled token-by-token from one 18B teacher and share one embedding space: a corpus indexed with 9B can be queried with 0.6B, lifting ViDoRe v3 from 62.3% to 63.5% at no added query cost.
More from Models
- d1-3B runs on a MacBook: hands-on demos show Liquid AI's open model in action — JosephJacks_ · 2026-10-08
- SentenceTransformers gets native ColPali model support thanks to tomaarsen — tomaarsen · 2026-10-08
- ChatGPT Plus users report 'thinking' time doubled in 2025 with no quality gain — Gazialp · 2026-10-08
- Mistral Large 4 debuts at #45 on Code Arena WebDev, near Opus 4.8 at 1/6 the price — arena · 2026-10-08
- Liquid AI releases Open d1: open-weight 3B and 600M multimodal decision models — JosephJacks_ · 2026-10-08
- Liquid AI details d1-omni-600M: 600M params for text+image or text+audio — JosephJacks_ · 2026-10-08