FULL STORY

EmbeddingGemma 2: From Launch to Hands-On Demos

Google released EmbeddingGemma 2, its first natively multimodal embedding model, on Oct 7, followed by Unsloth's quantized versions and quick community demos of in-browser local running and multimodal search.

2026-10-06 ~ 2026-10-07 · 4 episodes · 58 posts

Episode 1 · Unsloth Releases EmbeddingGemma-2 Multimodal Embedding Model (2026-10-06, 2 posts)

Unsloth released EmbeddingGemma-2, a multimodal embedding model supporting image, audio, and video feature extraction, along with a GGUF quantized version that topped the Hugging Face trending charts.

Episode 2 · Google Open-Sources EmbeddingGemma 2, First Natively Multimodal Embedding Model (2026-10-06, 52 posts)

Google CEO Sundar Pichai and Google DeepMind officially released EmbeddingGemma 2 around October 7. It is Google's first natively multimodal open-source embedding model and one of the few lightweight multimodal embedding solutions designed for on-device deployment, making it worth attention.

Confirmed

  • Official positioning: EmbeddingGemma 2 is DeepMind's first natively multimodal, open-source embedding model designed for on-device efficiency, going beyond text-only embeddings.
  • Scale: With only 740M parameters, it uses a modular encoder to map text (including code), images, video, and audio into a single embedding space (768 dimensions), and supports Matryoshka Representation Learning.
  • Openness: Open-sourced under the Apache 2.0 license, with weights available on Hugging Face and Kaggle; unsloth has already provided a GGUF version.
  • Use cases: Primarily aimed at offline, privacy-first RAG for on-device deployment.

Why it matters

  • Multimodal embedding models have previously relied on large cloud-based models or separate single-modality models. EmbeddingGemma 2 unifies vector representations across five modalities in a single 740M model, dramatically lowering the barrier to on-device retrieval and RAG deployment.
  • The Apache 2.0 license, combined with quick availability on Hugging Face/Kaggle and community GGUF versions, means developers can try it on local devices right away, accelerating privacy-first offline AI applications.

32 more related posts →

Episode 3 · Google EmbeddingGemma 2 Runs Locally in Browsers via WebGPU (2026-10-07, 2 posts)

After Google released EmbeddingGemma 2, its WebAI team's Jason Mayes shared a JS port demo that runs the embedding model fully client-side in browsers via WebGPU, with no server required.

Episode 4 · EmbeddingGemma 2 Demo Searches Images, Videos and Audio at Once (2026-10-07, 2 posts)

Developer tomaarsen showcased a multimodal embedding search demo powered by EmbeddingGemma 2, where a text query like 'Birds' retrieves related images, videos, and audio at once, with reverse search by image or sound also supported.