Google's multimodal embedding: vision token budget tunable from 70 to 1,120 per frame
tomaarsen · x · 2026-10-07
Details on Google's new multimodal embedding model: video defaults to 1 sampled frame per second, and audio should be mono at 16 kHz. The vision budget is configurable from 70 to 1,120 soft tokens per image/frame—more tokens buy finer visual detail at the cost of latency and context capacity. All modalities share an 8,192-token context window.
Related event: Google releases open-source multimodal EmbeddingGemma 2(27 posts)→
More from Multimodal
- Voice Agents Live or Die on 'Sounding Right': Turbo Shifts Tone With User Emotion — SucceededMind · 2026-10-07
- Magnific Original Series The Chronicles of Bone drops Chapter Six, made entirely with AI tools — Kavanthekid · 2026-10-07
- Hedra Lands in ChatGPT: Attach One Product Photo, Get a Full Commercial Ad — henloitsjoyce · 2026-10-07
- Marc Andreessen boosts AI film contest SLOPTOBERFEST grand prize to $25,000 — zealcaiden · 2026-10-07
- Image generation pricing leak: $0.05 per 2K image, $0.076 per 4K — op7418 · 2026-10-07
- Live human votes plugged into Flow-GRPO to stop image models gaming reward models — lmoroney · 2026-10-07