ColQwen-Omni retrieves audio and video zero-shot, no transcription needed
tomaarsen · x · 2026-08-18
Late interaction is expanding beyond page images: ColQwen-Omni takes text, images, audio and video — retrieving a recorded conversation is the same two calls, zero-shot, with no transcription step; the query "nausea" matches "carsickness" in the audio. ColPali-family checkpoints already match text queries against page images, charts and tables without OCR.
Related event: Sentence Transformers v6.0 ships with first-class late interaction models(33 posts)→
More from coding & agent
- PHAROS: An Open-Source npm for MCP Servers, Written in Go — Nofear001 · 2026-08-19
- Using AI Agents to draft release reports from evidence collections — CodeByPoonam · 2026-08-19
- MUON: An Open-Source Shared Brain for Parallel Coding Agents — Virtual_Gift_5327 · 2026-08-19
- Netlify integrates OpenRouter to enable model swapping without code changes — thisiskp_ · 2026-08-19
- Dev runs three Codex accounts plus Claude to parallelize coding agents — ChanceKelch · 2026-08-19
- DeepSeek open sources 'deepseek-harness' agent framework with 130k+ stars — alex_verem · 2026-08-19