pplx-embed-v2-late built on Qwen3.5, keeps one 128d vector per token with MaxSim
tomaarsen · x · 2026-10-08
Technical details of pplx-embed-v2-late: built on Qwen3.5 with bidirectional attention, keeping one 128d vector per token. Retrieval uses MaxSim late interaction — each query token takes its best match score in the document, then scores are summed, so different parts of a query can match different parts of a page.
Natively supported in Sentence Transformers via model.similarity().
Related event: Perplexity Open-Sources Multimodal Embedding Models pplx-embed-v2-late(38 posts)→
More from Models
- ChatGPT Work mode vs Codex: same quota, far more tasks done per 5-hour window — sasik520 · 2026-10-08
- Claude Haiku 5.5 ships with huge jumps: OSWorld 15.7%→72.4%, beats GPT-6 Luna across the board — mark_k · 2026-10-08
- OpenAI on GPT-6 Intelligent UI: the hard part is knowing when to show it — btibor91 · 2026-10-08
- Burkov slams watermarking in paid LLM outputs: 'I paid for this text' — burkov · 2026-10-08
- Epoch's InnovationEval: AI agents still far from producing real research innovations — Afinetheorem · 2026-10-08
- 113 decision models in 3 weeks: 70 built on Qwen, sub-cent per call — jonathanmalkin · 2026-10-08