EmbeddingGemma 2 Maps Text, Image, Video & Audio Into One Shared 768-Dim Space, 740M Params
tomaarsen · x · 2026-10-07
Sentence Transformers maintainer tomaarsen breaks down Google DeepMind's EmbeddingGemma 2: text (including code), images, video, and audio mapped into a single shared 768-dimensional embedding space; 740M total parameters, 100+ languages, Apache 2.0 license, with Sentence Transformers support. Thread with highlights.
Related event: Google releases open-source multimodal EmbeddingGemma 2(27 posts)→
More from Models
- Google's multimodal embeddinggemma-2 trends on Hugging Face — google · 2026-10-07
- Mistral CEO: Large 4 trained on our own compute, 'RL shows no sign of saturation' — sivareddyg · 2026-10-07
- Tesla's Grok voice goes hoarse but can't hear itself — a look at AI engineering shortcuts — PTrubey · 2026-10-07
- Perplexity ships open-weights pplx-decider-v1.1-27b at half the cost of v1 — perplexity_ai · 2026-10-07
- Mistral launches Large 4: 1T-param multimodal model, 49B active, open weights in October — beffjezos · 2026-10-07
- Perplexity's open-weights pplx-decider-v1.1-27b tops Hugging Face Decision Index 0.3 — AravSrinivas · 2026-10-07