0.6B visual embedding model rivals 8B models, keeps natural image search strong
antoine_chaffin · x · 2026-10-08
The author shares retrieval benchmark results for their visual embedding models. On ViDoRe (visual document retrieval), their 0.6B model is competitive with much larger models, close to Qwen3-VL-Embed-8B on MIRACL-Vision, while their 9B outperforms all compared models except Gemini-Embedding-2.\n\nOn PPLX-Q2I, an internal benchmark covering visual documents and natural images from production logs, both models beat Qwen3-VL-Embedding-8B by a wide margin, with the 9B close behind Gemini-Embedding-2. The author notes that although document retrieval is the primary use case, strong document retrieval performance does not come at the cost of natural image search.
Related event: Perplexity open-sources pplx-embed-v2-late retrieval models(36 posts)→
More from Models
- Anthropic Launches Claude Haiku 5.5, Costing ~75% Less Than Its Predecessor — rickasaurus · 2026-10-08
- Anthropic researcher's haiku hints Claude Haiku is jumping from 4.5 to 5.5 — edwinarbus · 2026-10-08
- Claude Haiku 5.5 spotted live in user accounts as rollout begins — rickasaurus · 2026-10-08
- Math professor grades OpenAI's 722 results: mostly B/C level, one D-level shock — khademinori · 2026-10-08
- Musk: Grok Bot requests will run on lightning-fast Grok 4.8 optimized for speed — XFreeze · 2026-10-08
- GitHub HydraFusion adds local model routing as Microsoft ships MAI Code 1.1 at 3-bit, 256K — BenBajarin · 2026-10-08