Google launches EmbeddingGemma 2, an open 740M multimodal embedding model for on-device use

Recoil42 · reddit · 2026-10-07

Google DeepMind released EmbeddingGemma 2, an open multimodal embedding model that maps text (including code), images, video, and audio—individually or combined—into a single 768-dimensional vector space. The 740M-parameter model combines a 270M text model with modular vision (170M) and audio (300M) encoders. It's designed to run on consumer hardware like phones and laptops, targeting low-latency on-device search, RAG, classification, and clustering.

Related event: Google Open-Sources EmbeddingGemma 2, First Native Multimodal Embedding Model(37 posts)→

Original post →

More from Models

Models channel →