OmniSearch puts text, images, audio, and video into one semantic search space

victorialslocum · x · 2026-07-21

Why unified multimodal search matters

OmniSearch demonstrates a shared semantic space for text, images, audio, and video, so a text query can retrieve audio or video and an image can find related content across modalities.

Key idea

Instead of converting everything into text first, the system embeds raw media directly:

Trade-offs

The post also notes the practical costs: audio and video need chunking, large vectors get expensive fast, and text-only corpora are still cheaper to handle with text embeddings.

Demo and notebooks are linked in the original post.

Original post →

More from Multimodal

Multimodal channel →