Google's multimodal embedding: vision token budget tunable from 70 to 1,120 per frame

tomaarsen · x · 2026-10-07

Details on Google's new multimodal embedding model: video defaults to 1 sampled frame per second, and audio should be mono at 16 kHz. The vision budget is configurable from 70 to 1,120 soft tokens per image/frame—more tokens buy finer visual detail at the cost of latency and context capacity. All modalities share an 8,192-token context window.

Related event: Google releases open-source multimodal EmbeddingGemma 2(27 posts)→

Original post →

More from Multimodal

Multimodal channel →