18B teacher distilled into 9B/0.6B students sharing one multi-vector embedding space

antoine_chaffin · x · 2026-10-08

A key feature: both models share the same embedding space, enabling cross-model retrieval — a convenient production setup. Training: an 18B teacher is trained contrastively, then distilled into 9B and 0.6B students via LEAF-style representation distillation aligning student outputs with the teacher per token, so both live in the same multi-vector space.

Why multimodal: real corpora aren't clean text — PDFs, slides, scans, charts, tables. Parsing + OCR is expensive and lossy; these models embed the rendered page directly, so text queries search the page itself, and natural images work too.

Related event: Perplexity open-sources pplx-embed-v2-late retrieval models(36 posts)→

Original post →

More from Models

Models channel →