Contextual AI releases 9B and 0.6B late-interaction embedding models hitting frontier retrieval results
antoine_chaffin · x · 2026-10-08
Contextual AI's antoinechaffin published a blog and model release: the in-house 9B ColBERT-style contextual model is now out, alongside a small 0.6B sibling, both achieving frontier results in text, multimodal and agentic retrieval.
He explains the switch to late interaction: dense models squeeze a whole document into one vector, which degrades with long/diverse documents and images, and a single dot product has limited expressivity ("hi LIMIT"). One vector per token + MaxSim lifts both bottlenecks.
Related event: Perplexity Open-Sources Multimodal Embedding Models pplx-embed-v2-late(38 posts)→
More from Models
- ChatGPT on Browser Shows 'Capabilities Reduced' Warning — A New Kind of Rate Limit? — jasondeanlee · 2026-10-08
- ChatGPT Work mode vs Codex: same quota, far more tasks done per 5-hour window — sasik520 · 2026-10-08
- Epoch's InnovationEval: AI agents still far from producing real research innovations — Afinetheorem · 2026-10-08
- 113 decision models in 3 weeks: 70 built on Qwen, sub-cent per call — jonathanmalkin · 2026-10-08
- User reports Haiku 5.5 is a major workflow upgrade in screenshot post — Sorcerer12345 · 2026-10-08
- OpenRouter launches Decision Model Rankings, with typesafeai leading all categories — gaganghotra_ · 2026-10-08