Google ships EmbeddingGemma 2: 740M open model embedding text, image, video and audio
tomaarsen · x · 2026-10-07
Google DeepMind released EmbeddingGemma 2, an Apache 2.0 multimodal embedding model that maps text (including code), images, video and audio into one shared 768-dimensional vector space, so a text query can retrieve photos, audio clips or videos.
- 740M total params, 100+ languages, first-class Sentence Transformers support (model.encode() / model.similarity())
- Modular footprint: text 270M, vision 170M, audio 300M; load only needed encoders (text-only 270M, text+image 440M, text+audio 570M), designed for phones and laptops
- Vs. v1: MTEB multilingual v2 61.36 vs 61.15; code retrieval jumps from 68.76 to 78.68 on MTEB code v1 (14% relative)
- Multimodal scores at 768d: MMEB v2 visual documents 67.84 NDCG@5, video 50.67 Hit@1, MSEB audio retrieval 69.54 MRR@10
- Matryoshka-trained 768/512/256/128 dims: 256d cuts storage to a third with MTEB multilingual 61.36→60.41; 128d drops notably (MMEB v2 59.01→45.65), better reserved for text-only; mixed modalities fine at 256d
- Warning: use BF16 or FP32, never FP16 — activations exceed FP16's dynamic range, causing NaNs or silently degraded embeddings
More from Infra
- Qwen3.8 Flash Next IQ1_M hits 55 tok/s on a 5060 Ti 16GB and still codes well — bobaburger · 2026-10-07
- EmbeddingGemma 2 runs offline multimodal RAG on phones in ~191MB–567MB of RAM — Saboo_Shubham_ · 2026-10-07
- Qwen 3.8 27B hits 96 t/s decode with 110k context on a single 16GB RTX 5080 via NInfer — Kernoriordan · 2026-10-07
- OpenAI to fund Nicholas Nethercote's work speeding up the Rust compiler — charliermarsh · 2026-10-07
- Google signs 20-year nuclear PPA with Constellation: 3,590 MW and $4.3B investment — demian_ai · 2026-10-07
- NVIDIA's Vera CPU spins up 2,000 sandboxes in 27.4s, 8x faster than rivals — mattturck · 2026-10-07