ColQwen-Omni retrieves audio and video zero-shot, no transcription needed

tomaarsen · x · 2026-08-18

Late interaction is expanding beyond page images: ColQwen-Omni takes text, images, audio and video — retrieving a recorded conversation is the same two calls, zero-shot, with no transcription step; the query "nausea" matches "carsickness" in the audio. ColPali-family checkpoints already match text queries against page images, charts and tables without OCR.

Related event: Sentence Transformers v6.0 ships with first-class late interaction models(33 posts)→

Original post →

More from coding & agent

coding & agent channel →