Google releases EmbeddingGemma 2: a 740M Apache 2.0 multimodal embedding model for fully local semantic search
lmoroney · x · 2026-10-07
Google has released EmbeddingGemma 2, a 740M-parameter embedding model built on Gemma 4 and licensed under Apache 2.0. It maps text, code, images, audio, and video into a single shared embedding space, so a text query can retrieve a video clip or voice memo entirely on-device.
Key details:
- Modular design: text-only work runs on as little as 270M parameters, with optional vision (170M) and audio (300M) encoders
- Vectors can be truncated from 768 to 512/256/128 dimensions
- Quantized, it needs roughly 191MB of active RAM for text and 567MB for the full multimodal version on a Pixel 11 Pro
- Context window is now 8K tokens, 4x the original
The author suggests running your own retrieval-quality comparison at 768 vs 256 dimensions on a sample of your documents before choosing a setting for local RAG.
More from Infra
- NVIDIA's NeMo-DCR Cuts Trillion-Parameter RL Weight Sync from 87.5 min to 150s — nvidia · 2026-10-07
- jiti-lfe replicates across 9 hosts in 13 minutes — arthurcolle · 2026-10-07
- Qwen3.8 Flash Next GGUF benchmark: IQ3_S the sweet spot, 42.7M tokens tested — lxfater · 2026-10-07
- Used PS5 Pro hits $1,399 at GameStop as AI datacenters squeeze memory supply — aakashgupta · 2026-10-07
- Musk: xAI will build and run its Terafab itself, TSMC may only sublease part of it — MickeySteamboat · 2026-10-07
- "72% of the intelligence with 3.8% of the GPUs": Mistral's compute-efficiency ratio sparks debate — cyb3rops · 2026-10-07