Cross-model retrieval tested: 0.6B querying a 9B index closes half the gap at no latency cost
antoine_chaffin · x · 2026-10-08
This post highlights cross-model retrieval enabled by the shared embedding space: using the 0.6B model to query a 9B index beats the 0.6B index with no query-time cost. The 9B for querying remains strongest, but the smaller model closes roughly half the gap — a flexible accuracy/latency tradeoff for production.
Related event: Perplexity Open-Sources Multimodal Embedding Models pplx-embed-v2-late(38 posts)→
More from Infra
- Tracking a 24/7 agent for 30 days: $6 VPS, $22 API, and uptime is the real leak — YamOk7317 · 2026-10-08
- Bain: nearly 150GW of new data center capacity by 2030, requiring $5-6.5 trillion — Beth_Kindig · 2026-10-08
- TypeSafeAI's Jev served a trillion tokens on Modal within three days of launch — josh_wills · 2026-10-08
- llama.cpp PR adds GPU cache for host-resident MoE experts, big speedup potential — jacek2023 · 2026-10-08
- Mapping the firm-power stack behind Google's massive nuclear deal: CEG, TLN, VST, Hubbell, Bloom — demian_ai · 2026-10-08
- Why temp=0 LLM inference still isn't deterministic: floating-point order and parallelism — ducha_aiki · 2026-10-08