EmbeddingGemma 2 Maps Text, Image, Video & Audio Into One Shared 768-Dim Space, 740M Params

tomaarsen · x · 2026-10-07

Sentence Transformers maintainer tomaarsen breaks down Google DeepMind's EmbeddingGemma 2: text (including code), images, video, and audio mapped into a single shared 768-dimensional embedding space; 740M total parameters, 100+ languages, Apache 2.0 license, with Sentence Transformers support. Thread with highlights.

Related event: Google releases open-source multimodal EmbeddingGemma 2(27 posts)→

Original post →

More from Models

Models channel →