Cross-model retrieval: pplx-embed-v2's 0.6B querying a 9B index beats a 0.6B index
antoine_chaffin · x · 2026-10-08
A key feature of the pplx-embed-v2 release: the 9B and 0.6B models share the same embedding space, enabling cross-model retrieval.
Perplexity's team benchmarked querying a 9B index with the 0.6B model against symmetric 9B and 0.6B setups. The smaller model querying the 9B index scores higher than using a 0.6B index — an asymmetric production option that pairs cheap queries with a high-quality index.
More from Infra
- Windows demo routes coding tasks to local model with GPU spinning, llama.cpp lands on Windows ML — ryanshrout · 2026-10-08
- exe.dev explains crossing the hyper-thread boundary: core scheduling cookies for VM isolation — davidcrawshaw · 2026-10-08
- NVIDIA's LoGRA cuts RL training memory by up to 45.7%, trains 27B model where Adam OOMs — mark_k · 2026-10-08
- Qwen3.8-Flash-Next on 6x3090 without NVLink: prefill 8-10x faster, long-context decode 2-3x — flynth92 · 2026-10-08
- Google DeepMind's Philipp Schmid: Give Every AI Agent Its Own Managed Cloud Sandbox — AI Engineer · 2026-10-08
- GitHub goes down; engineers confirm a fix is in progress — iannuttall · 2026-10-08