Google's multimodal embedding: images cost 280 tokens, audio 25 tokens per second
tomaarsen · x · 2026-10-07
Google's new multimodal embedding model gives all modalities a shared 8,192-token context window. Default costs: 280 tokens per image, 140 per video frame, and 25 per second of audio—about 327 seconds of audio-only input. Mixed inputs draw from the same budget, and media markers let you embed a product listing (text, photos, demo video) as one vector searchable by a text query.
Related event: Google releases open-source multimodal EmbeddingGemma 2(27 posts)→
More from Multimodal
- Voice Agents Live or Die on 'Sounding Right': Turbo Shifts Tone With User Emotion — SucceededMind · 2026-10-07
- Magnific Original Series The Chronicles of Bone drops Chapter Six, made entirely with AI tools — Kavanthekid · 2026-10-07
- Hedra Lands in ChatGPT: Attach One Product Photo, Get a Full Commercial Ad — henloitsjoyce · 2026-10-07
- Marc Andreessen boosts AI film contest SLOPTOBERFEST grand prize to $25,000 — zealcaiden · 2026-10-07
- Image generation pricing leak: $0.05 per 2K image, $0.076 per 4K — op7418 · 2026-10-07
- Live human votes plugged into Flow-GRPO to stop image models gaming reward models — lmoroney · 2026-10-07