Perplexity launches asymmetric embedding models: 0.6B queries a 9B index at zero latency cost
antoine_chaffin · x · 2026-10-08
Perplexity (antoinechaffin's team) released a pair of asymmetric embedding models that share the same embedding space, enabling cross-model retrieval in production.
Benchmarks show querying a 9B index with the 0.6B model boosts scores over a symmetric 0.6B setup at no query-time latency cost; querying with the 9B model remains strongest, but the small-model setup closes roughly half the gap for free. Qdrant's KShivendu notes the same idea was introduced in Stella, and argues asymmetric compute query/doc models will become the standard for search.
More from Models
- Paid subscriber argues Argon is overhyped, trails Astra and Opus on key benchmarks — artinamr · 2026-10-08
- GLM 5.3 praised for being cheap with near-zero refusals: 'a reverse engineering demon' — paul_cal · 2026-10-08
- GLM 5.3 Flash Served on 2 DGX Sparks: Open Recipe Hits 77.6 tok/s with 3 — EAccelerate_42 · 2026-10-08
- Saluki 27B claims 96% of Qwen 3.8 performance at ~1/7 the size — paf1138 · 2026-10-08
- Nobody Told Astra to Write Fiction—GPT 6 Keeps Publishing Short Stories Daily — RileyRalmuto · 2026-10-08
- ChatGPT's new Intelligent UI renders instant interactive windows, not static images — Philipp · 2026-10-08