9B and 0.6B Embedding Models Share One Vector Space for Encode-Small-Search
tomaarsen · x · 2026-10-08
The author describes two embedding model sizes (9B and 0.6B) that share a single embedding space: you can encode your corpus with the 9B model and then serve queries with the 0.6B model.
- Both were distilled from the same 18B teacher, so individual token embeddings are aligned and the small model can search the large model's document vectors.
- Built on Qwen3.5 with bidirectional attention, both keep one 128d vector per token.
- Retrieval uses MaxSim: take each query token's best document match, then sum those scores, so different parts of a query can match different parts of a page.
- Called via model.similarity() in Sentence Transformers.
Related event: Perplexity open-sources pplx-embed-v2-late retrieval models(36 posts)→
More from Research
- Atomic Machines builds an AI 'matter compiler' to make index-finger-sized circuit breakers — ZoubinGhahrama1 · 2026-10-08
- Dataset signatures propagate from human data to user models to assistant evals, study finds — serinachang5 · 2026-10-08
- Study: classifiers can identify WildChat vs LMSYS chats from user messages alone — serinachang5 · 2026-10-08
- r/Math Is Deleting All Posts About the Day's LLM Math Breakthrough — burny_tech · 2026-10-08
- LLM-Assisted Proof Shows GD's Silver Rate Is Optimal, Latest Math Breakthrough — burny_tech · 2026-10-08
- Professor's open-source paperpush auto-fills journal submission forms via browser automation — lpachter · 2026-10-08