Shared embedding space: index with 9B, serve queries with 0.6B
tomaarsen · x · 2026-10-08
The 0.6B and 9B share one embedding space: encode your corpus with 9B, serve queries with 0.6B — near-large performance at low cost. Both were distilled from the same 18B teacher by aligning individual token embeddings, so the small model can search the large model's document vectors. The thread notes the architecture: Qwen3.5-based bidirectional attention, one 128d vector per token, MaxSum scoring via per-query-token best matches, exposed as model.similarity() in Sentence Transformers.
Related event: Perplexity Open-Sources Multimodal Embedding Models pplx-embed-v2-late(38 posts)→
More from Models
- ChatGPT on Browser Shows 'Capabilities Reduced' Warning — A New Kind of Rate Limit? — jasondeanlee · 2026-10-08
- ChatGPT Work mode vs Codex: same quota, far more tasks done per 5-hour window — sasik520 · 2026-10-08
- Epoch's InnovationEval: AI agents still far from producing real research innovations — Afinetheorem · 2026-10-08
- 113 decision models in 3 weeks: 70 built on Qwen, sub-cent per call — jonathanmalkin · 2026-10-08
- User reports Haiku 5.5 is a major workflow upgrade in screenshot post — Sorcerer12345 · 2026-10-08
- OpenRouter launches Decision Model Rankings, with typesafeai leading all categories — gaganghotra_ · 2026-10-08