Perplexity's pplx-embed-v2-late: serve 0.6B queries against 9B document embeddings

tomaarsen · x · 2026-10-08

Perplexity released pplx-embed-v2-late, a pair of models (0.6B and 9B) for text and image retrieval. Key result: combining the 0.6B query model with the 9B document encoder beats 0.6B on both sides — text six-domain average nDCG@10 goes from 78.0 to 79.6, and ViDoRe v3 images from 62.3 to 63.5.

The two sizes share an embedding space: encode your corpus once with 9B, then serve queries with 0.6B latency. Both were distilled from the same 18B teacher with individual token embeddings aligned, so the small model can search the large model's document vectors directly.

Related event: Perplexity Open-Sources Multimodal Embedding Models pplx-embed-v2-late(38 posts)→

Original post →

More from Models

Models channel →