18B teacher distilled into 9B/0.6B students sharing one multi-vector embedding space
antoine_chaffin · x · 2026-10-08
A key feature: both models share the same embedding space, enabling cross-model retrieval — a convenient production setup. Training: an 18B teacher is trained contrastively, then distilled into 9B and 0.6B students via LEAF-style representation distillation aligning student outputs with the teacher per token, so both live in the same multi-vector space.
Why multimodal: real corpora aren't clean text — PDFs, slides, scans, charts, tables. Parsing + OCR is expensive and lossy; these models embed the rendered page directly, so text queries search the page itself, and natural images work too.
Related event: Perplexity open-sources pplx-embed-v2-late retrieval models(36 posts)→
More from Models
- Grok bots 0.68.1 add slide decks, formatted email, and 1920x1200 computer use — Daniel_Farinax · 2026-10-08
- GLM V4.1 Looks Like the Best Chinese Model on ARC-2, Says TeortaxesTex — teortaxesTex · 2026-10-08
- Models systematically underestimate their own capabilities, even newest ones — repligate · 2026-10-08
- Nace.AI open-sources Drex 1.1, an 8B diffusion-LM decision model with released weights — nischay_twt · 2026-10-08
- Indie chatbot Auro lets users blind-pick between model versions to shape its personality — TheMoonMidas · 2026-10-08
- Claude Haiku 5.5 spotted in Claude Code update, rumored at $0.1/M input tokens — kimmonismus · 2026-10-08