Google ships EmbeddingGemma 2: 740M multimodal embeddings that run on phones
dl_weekly · x · 2026-10-11
Google released EmbeddingGemma 2, an open multimodal embedding model extending the line beyond text to images, audio and video in one shared embedding space — small enough to run on a smartphone.
- Size: 740M parameters (2x v1), mostly vision/audio encoders that text-only apps can drop; the 270M text core uses 191MB quantized on a Pixel 11 Pro
- Compression: 768-dim embeddings can be truncated to 128 via Matryoshka Representation Learning (6x storage savings); at 256 dims retrieval keeps 95% quality
- Benchmarks: 78.68 on MTEB's code section, nearly 10 points above v1
- Traction: the original model has been downloaded 20M+ times
Example use case: matching a voice memo to the right moment in a video, entirely on-device.
More from Models
- Fireworks' Forge conference lands Jensen Huang, DHH and swyx for a day on open-model frontier — sophiamyang · 2026-10-11
- Hirundo Releases 'Westernized' Qwen Finetunes After Reddit Mods Delete Discussion — firstcenturyman · 2026-10-11
- Leaked: OpenAI's unreleased model solved most math problems in a single prompt, ~3h each — Puzzleheaded-King584 · 2026-10-11
- Turing Post's 2026 guide: which LLM benchmarks to use for reasoning, coding, math and agents — TheTuringPost · 2026-10-11
- OpenAI's unreleased model reportedly solved most math problems in a single prompt — Puzzleheaded-King584 · 2026-10-11
- Musk says RAM supply is the problem as quantized 25GB MoE local run speculated for 2028-29 — elonmusk · 2026-10-11