One text query retrieves photos, audio and video in EmbeddingGemma 2's shared space
tomaarsen · x · 2026-10-07
With EmbeddingGemma 2, a text query can retrieve photos, audio recordings or video clips — every modality becomes a dense vector in the same space. It keeps the familiar Sentence Transformers model.encode() and model.similarity() API, and the model card includes a short text-search quickstart.
More from Models
- llama.cpp ships Day-0 support for Google's EmbeddingGemma 2 — ggerganov · 2026-10-07
- Trying to Plug Open-Source Mistral Into an Agentic Coder Just Doesn't Work, Says Berman — MatthewBerman · 2026-10-07
- antirez: DeepSeek v4.1 outscores Mistral Large 4 on DeepSWE 1.1 and other benchmarks — antirez · 2026-10-07
- StartLux claims its 27B model beats Jev AI on 31 of 38 benchmarks, self-reported results — Dr_Singularity · 2026-10-07
- 24 models tested on 669 clinical decisions: Jev stays #1 as two free models close in — MaziyarPanahi · 2026-10-07
- Never run EmbeddingGemma 2 in FP16: use BF16 or FP32 to avoid NaNs — tomaarsen · 2026-10-07