Google releases EmbeddingGemma 2: 740M multimodal embeddings for on-device RAG
jacek2023 · reddit · 2026-10-06
Google DeepMind released EmbeddingGemma 2, an open multimodal embedding model now on Hugging Face (with a GGUF build from unsloth). It packs 740M total parameters and maps text (incl. code), images, video and audio into a single unified 768-dimensional vector space, targeting phones and laptops.
Key points:
- Architecture: 270M text backbone (130M transformer + 140M embedder) plus selectively loadable vision (170M) and audio (300M) encoders
- 100+ languages; 14% improvement on code tasks over its predecessor
- Native Matryoshka Representation Learning: truncatable to 128d/256d/512d/768d for up to 6x vector storage savings
- 8K context window, handling minutes of audio or video
- Task-steered embeddings via lightweight instruction prefixes for search, classification, clustering and similarity
Aimed at on-device search, RAG, classification and clustering.
More from Infra
- Agentic data toll: enterprise agent data services to hit ~$30B by 2030, says Bajarin — BenBajarin · 2026-10-07
- DGX Station + one RTX Pro 6K hits 63K tok/s prefill on DeepSeek V4.1 Flash — Sentdex · 2026-10-07
- Huawei reportedly testing 256K-card Atlas-950 SuperPoD aiming for million-card compute — teortaxesTex · 2026-10-07
- MIT CSAIL Launches Ascent Lab, a Browser Platform for Robot Learning, With Solana Token — MIT_CSAIL · 2026-10-07
- NVIDIA Livestreams Smart Routing for Hybrid AI on DGX Spark — NVIDIAAI · 2026-10-07
- One RTX Pro 6K Sidecar Lifts DGX Station GB300 to 60K tok/s Prefill with DeepSeek V4.1 Flash — Sentdex · 2026-10-07