Google open-sources EmbeddingGemma2: a 740M omni-modal embedding model that runs on phones
AGI Hunt · wechat · 2026-10-07
Google CEO Sundar Pichai and Google DeepMind jointly announced the open-source release of EmbeddingGemma2, billed as Google's first open-source natively multimodal embedding model. Weights are on Hugging Face under Apache 2.0.
Architecture: the 740M total combines a 270M text backbone (130M transformer + 140M embedder), a 170M vision encoder and a 300M audio encoder; the latter two are optional, giving four sizes — text-only 270M, text+image 440M, text+audio 570M, full 740M. It builds on the Gemma 4 architecture, sharing its text tokenizer and audio encoder.
Key specs: 768-dim output, 100+ languages, 8K context (4x the previous generation) fitting 5.5 minutes of audio, 29 images or 58 video frames at once; Matryoshka representation learning lets 768 dims truncate to 512/256/128, cutting vector-store storage up to 6x; text inputs accept task prefixes. Quantized, it uses 191MB RAM for text-only and 567MB for full modality on a Pixel 11 Pro.
Two pitfalls: truncating to 256 dims barely loses accuracy, but 128 dims drops multimodal scores sharply (MMEBv2 from 59.01 to 45.65) — use 128 only for text; don't run float16, since activations exceed its range and silently produce NaN or degraded results, so use bfloat16 or float32.
Scores (768-dim, full precision): MTEB multilingual v2 61.36 (vs 61.15 prior), MTEB code v1 78.68 (vs 68.76, up nearly 10 points); image MIEB lite 64.64, MMEBv2 Image 57.28, VisDoc 67.84, Video 50.67; audio MSEB retrieval 69.54, MAEB 49.39. Google positions it as one of the strongest sub-1B multimodal embedding models.
Vs. rivals: text-only trails Qwen3-Embedding-0.6B by 3 points and MMEB-V2 trails Qwen3-VL-Embedding-2B by 14 points, but the latter starts at 2B and lacks audio; OpenAI's text-embedding-3-large scores 58.93 on MTEB multilingual, beaten by EmbeddingGemma2's 270M text part by over 2 points, and OpenAI embeddings remain text-only.
Ecosystem and demos: supported by transformers, sentence-transformers, MLX, vLLM, llama.cpp, SGLang, Ollama and LM Studio, with browser support via transformers.js or WebGPU and a fine-tuning tutorial from Unsloth. Four official demos: on-device album search (InstantMediaSearch), video moment finder, local file retrieval (Foresight), and a real-time MediaPipe Decision Task API demo where the 270M model decides in 43ms per step versus 640ms for a regular LLM.
Context: the previous EmbeddingGemma (308M, text-only, Gemma license) shipped last September and has passed 20M downloads; this release comes 13 months later under Apache 2.0. The piece also recounts OpenAI's April mistaken deprecation notice for text-embedding-3-small, which sparked developer debate about the risk of closed embedding models going offline.
More from Models
- Dev warns OpenRouter share, cache hit rate, latency stats are easily gamed for marketing — charles_irl · 2026-10-07
- OpenAI's unreleased model reportedly proves quasi-Riemann hypothesis with Lean proof — ChrisGPT · 2026-10-07
- Fed Claude Opus my blurry handheld Saturn shots, it fused them into one best image — adonis_singh · 2026-10-07
- Two Labs, One Race: Anthropic vs OpenAI Frontier Model Release Timeline, 2023–2026 — Medical-Sky7620 · 2026-10-07
- Claude was given robot skin — Opus was curious but anxious about hooking up — repligate · 2026-10-07
- Gemini 2.5 Pro retiring October 20, 2026, users say goodbye — hargup13 · 2026-10-07