FULL STORY
EmbeddingGemma 2: From Launch to Hands-On Demos
Google released EmbeddingGemma 2, its first natively multimodal embedding model, on Oct 7, followed by Unsloth's quantized versions and quick community demos of in-browser local running and multimodal search.
2026-10-06 ~ 2026-10-07 · 4 episodes · 58 posts
Episode 1 · Unsloth Releases EmbeddingGemma-2 Multimodal Embedding Model (2026-10-06, 2 posts)
Unsloth released EmbeddingGemma-2, a multimodal embedding model supporting image, audio, and video feature extraction, along with a GGUF quantized version that topped the Hugging Face trending charts.
- Unsloth ships EmbeddingGemma-2, a multimodal embedding model with vision and audio — unsloth · 2026-10-06
- unsloth's EmbeddingGemma-2 GGUF quantization trends on Hugging Face — unsloth · 2026-10-07
Episode 2 · Google Open-Sources EmbeddingGemma 2, First Natively Multimodal Embedding Model (2026-10-06, 52 posts)
Google CEO Sundar Pichai and Google DeepMind officially released EmbeddingGemma 2 around October 7. It is Google's first natively multimodal open-source embedding model and one of the few lightweight multimodal embedding solutions designed for on-device deployment, making it worth attention.
Confirmed
- Official positioning: EmbeddingGemma 2 is DeepMind's first natively multimodal, open-source embedding model designed for on-device efficiency, going beyond text-only embeddings.
- Scale: With only 740M parameters, it uses a modular encoder to map text (including code), images, video, and audio into a single embedding space (768 dimensions), and supports Matryoshka Representation Learning.
- Openness: Open-sourced under the Apache 2.0 license, with weights available on Hugging Face and Kaggle; unsloth has already provided a GGUF version.
- Use cases: Primarily aimed at offline, privacy-first RAG for on-device deployment.
Why it matters
- Multimodal embedding models have previously relied on large cloud-based models or separate single-modality models. EmbeddingGemma 2 unifies vector representations across five modalities in a single 740M model, dramatically lowering the barrier to on-device retrieval and RAG deployment.
- The Apache 2.0 license, combined with quick availability on Hugging Face/Kaggle and community GGUF versions, means developers can try it on local devices right away, accelerating privacy-first offline AI applications.
- Google releases EmbeddingGemma 2: 740M multimodal embeddings for on-device RAG — jacek2023 · 2026-10-06
- Google launches EmbeddingGemma 2, a 740M open multimodal embedding model for on-device use — sundarpichai · 2026-10-07
- Google DeepMind details EmbeddingGemma 2: unified embeddings for code, image, audio, video — GoogleDeepMind · 2026-10-07
- Google releases EmbeddingGemma 2: a 740M-parameter natively multimodal open embedding model — GoogleDeepMind · 2026-10-07
- Google ships EmbeddingGemma 2: 740M multimodal embeddings under Apache 2.0 — GlennCameronjr · 2026-10-07
- Google releases EmbeddingGemma 2: open multimodal embedding model from 270M to 740M params — osanseviero · 2026-10-07
- Google releases EmbeddingGemma 2: open 270m-740m multimodal embeddings for on-device use — osanseviero · 2026-10-07
- Google Releases EmbeddingGemma 2: 740M Open Multimodal Embedding Model Running on 0.5GB RAM — gnukeith · 2026-10-07
- EmbeddingGemma 2 Maps Text, Image, Video & Audio Into One Shared 768-Dim Space, 740M Params — tomaarsen · 2026-10-07
- Google ships EmbeddingGemma 2: 740M open model embedding text, image, video and audio — tomaarsen · 2026-10-07
- One text query retrieves photos, audio and video in EmbeddingGemma 2's shared space — tomaarsen · 2026-10-07
- EmbeddingGemma 2 is modular: load only 270M for text-only on phones and laptops — tomaarsen · 2026-10-07
- EmbeddingGemma 2 code retrieval jumps 14%, multilingual gains marginal over v1 — tomaarsen · 2026-10-07
- EmbeddingGemma 2 multimodal scores: 67.84 visual docs, 50.67 video, 69.54 audio — tomaarsen · 2026-10-07
- EmbeddingGemma 2 dim tradeoff: 128d shrinks vectors 6x but MMEB drops to 45.65 — tomaarsen · 2026-10-07
- Matryoshka training lets embeddings truncate to 256d with a third of the storage — tomaarsen · 2026-10-07
- Task prefixes for embedding: pair SearchQuery with Document for retrieval — tomaarsen · 2026-10-07
- One vector for a whole product listing: multimodal embedding mixes text, photos and video — tomaarsen · 2026-10-07
- Google's multimodal embedding: images cost 280 tokens, audio 25 tokens per second — tomaarsen · 2026-10-07
- Google's multimodal embedding: vision token budget tunable from 70 to 1,120 per frame — tomaarsen · 2026-10-07
Episode 3 · Google EmbeddingGemma 2 Runs Locally in Browsers via WebGPU (2026-10-07, 2 posts)
After Google released EmbeddingGemma 2, its WebAI team's Jason Mayes shared a JS port demo that runs the embedding model fully client-side in browsers via WebGPU, with no server required.
- EmbeddingGemma 2 runs locally in-browser via WebGPU with live demo — xenovatech · 2026-10-07
- Google's EmbeddingGemma2 gets a client-side WebAI JavaScript port you can run free in-browser — jason_mayes · 2026-10-07
Episode 4 · EmbeddingGemma 2 Demo Searches Images, Videos and Audio at Once (2026-10-07, 2 posts)
Developer tomaarsen showcased a multimodal embedding search demo powered by EmbeddingGemma 2, where a text query like 'Birds' retrieves related images, videos, and audio at once, with reverse search by image or sound also supported.
- EmbeddingGemma 2 Demo: One Search Across Images, Video and Audio in Browser — tomaarsen · 2026-10-07
- EmbeddingGemma 2 demo: search images, videos and audio with one text query — tomaarsen · 2026-10-07