SentenceTransformers V6 Unifies Text, Vision, and Audio in Single API

ManuelFaysse · x · 2026-08-21

SentenceTransformers V6 is released with a unified API that supports ColPali, ColQwen, and other visual retrieval models out of the box. This update enables training and inference for text, vision, and audio within a single framework. The standalone ColPali repository will be deprecated in favor of the main library, ensuring long-term support from the Hugging Face team.

Original post →

More from Models

Models channel →