How EmbeddingGemma 2 works: six steps from five modalities to one 768-dim vector
KyeGomezB · x · 2026-10-10
Google's EmbeddingGemma 2 maps text, code, images, video, and audio into a shared vector space for multimodal search and retrieval. The pipeline has six steps:
- Text encoding: 24-layer Transformer with local and global attention.
- Vision encoding: 16-layer ViT converts images into visual tokens.
- Audio encoding: 12-layer Conformer turns audio features into audio tokens.
- Multimodal fusion: visual and audio tokens are aligned with text tokens and processed jointly.
- Projection: pooled into a single 768-dim, L2-normalized vector.
- Matryoshka embeddings: reducible to 512/256/128 dims to save storage and compute.
More from Models
- StepFun's Step 5 Preview free in Kilo for a week: 600B params, 27B active, 1M-token context — StepFun_ai · 2026-10-10
- OpenAI reportedly solved 92 of the 500 most important open math problems in one GitHub push — altryne · 2026-10-10
- Polymarket odds: only 55% chance xAI ships Grok 5 by end of 2026 — Polymarket · 2026-10-10
- After index bug fixes, gpt-live-1 tops Artificial Analysis speech-to-speech ranking — pbbakkum · 2026-10-10
- DeepSeek-V4 answers flip with 2-token input shifts; NIAH swings 40 points — Francis_YAO_ · 2026-10-10
- Artificial Analysis teases AA-Robotics: frontier models zero-shot robot arm control — ArtificialAnlys · 2026-10-10