Trained on 186M pairs from 594 datasets in 46 languages, with benchmark data fully removed

antoine_chaffin · x · 2026-10-08

Training details: 186M query-document pairs from 594 datasets in 46 languages, mixing text-to-text, text-to-image and text-to-visual-document. Every dataset associated with the evaluated benchmarks was removed — costing some good data, but the authors consider trustworthy numbers worth it.

Related event: Perplexity open-sources pplx-embed-v2-late retrieval models(36 posts)→

Original post →

More from Models

Models channel →