Shared embedding space: index with 9B, serve queries with 0.6B

tomaarsen · x · 2026-10-08

The 0.6B and 9B share one embedding space: encode your corpus with 9B, serve queries with 0.6B — near-large performance at low cost. Both were distilled from the same 18B teacher by aligning individual token embeddings, so the small model can search the large model's document vectors. The thread notes the architecture: Qwen3.5-based bidirectional attention, one 128d vector per token, MaxSum scoring via per-query-token best matches, exposed as model.similarity() in Sentence Transformers.

Related event: Perplexity Open-Sources Multimodal Embedding Models pplx-embed-v2-late(38 posts)→

Original post →

More from Models

Models channel →