Trained on 186M pairs from 594 datasets in 46 languages, with benchmark data fully removed
antoine_chaffin · x · 2026-10-08
Training details: 186M query-document pairs from 594 datasets in 46 languages, mixing text-to-text, text-to-image and text-to-visual-document. Every dataset associated with the evaluated benchmarks was removed — costing some good data, but the authors consider trustworthy numbers worth it.
Related event: Perplexity open-sources pplx-embed-v2-late retrieval models(36 posts)→
More from Models
- GLM V4.1 Looks Like the Best Chinese Model on ARC-2, Says TeortaxesTex — teortaxesTex · 2026-10-08
- Models systematically underestimate their own capabilities, even newest ones — repligate · 2026-10-08
- Nace.AI open-sources Drex 1.1, an 8B diffusion-LM decision model with released weights — nischay_twt · 2026-10-08
- Indie chatbot Auro lets users blind-pick between model versions to shape its personality — TheMoonMidas · 2026-10-08
- Claude Haiku 5.5 spotted in Claude Code update, rumored at $0.1/M input tokens — kimmonismus · 2026-10-08
- Haiku 5.5 day: Anthropic's latest small model appears to ship — scaling01 · 2026-10-08