Google's multimodal embedding: images cost 280 tokens, audio 25 tokens per second

tomaarsen · x · 2026-10-07

Google's new multimodal embedding model gives all modalities a shared 8,192-token context window. Default costs: 280 tokens per image, 140 per video frame, and 25 per second of audio—about 327 seconds of audio-only input. Mixed inputs draw from the same budget, and media markers let you embed a product listing (text, photos, demo video) as one vector searchable by a text query.

Related event: Google releases open-source multimodal EmbeddingGemma 2(27 posts)→

Original post →

More from Multimodal

Multimodal channel →