Perplexity's pplx-embed-v2-late: serve 0.6B queries against 9B document embeddings
tomaarsen · x · 2026-10-08
Perplexity released pplx-embed-v2-late, a pair of models (0.6B and 9B) for text and image retrieval. Key result: combining the 0.6B query model with the 9B document encoder beats 0.6B on both sides — text six-domain average nDCG@10 goes from 78.0 to 79.6, and ViDoRe v3 images from 62.3 to 63.5.
The two sizes share an embedding space: encode your corpus once with 9B, then serve queries with 0.6B latency. Both were distilled from the same 18B teacher with individual token embeddings aligned, so the small model can search the large model's document vectors directly.
Related event: Perplexity Open-Sources Multimodal Embedding Models pplx-embed-v2-late(38 posts)→
More from Models
- ChatGPT Work mode vs Codex: same quota, far more tasks done per 5-hour window — sasik520 · 2026-10-08
- Claude Haiku 5.5 ships with huge jumps: OSWorld 15.7%→72.4%, beats GPT-6 Luna across the board — mark_k · 2026-10-08
- OpenAI on GPT-6 Intelligent UI: the hard part is knowing when to show it — btibor91 · 2026-10-08
- Burkov slams watermarking in paid LLM outputs: 'I paid for this text' — burkov · 2026-10-08
- Epoch's InnovationEval: AI agents still far from producing real research innovations — Afinetheorem · 2026-10-08
- 113 decision models in 3 weeks: 70 built on Qwen, sub-cent per call — jonathanmalkin · 2026-10-08