Google DeepMind on speech-to-speech: conversational, intelligent, multimodal — pick two
AI Engineer · youtube · 2026-09-15
Google DeepMind's Valeria Wu Fon and Tom Ouyang explain their speech-to-speech work in Gemini:
- Pre-2018 speech recognition was a cascade of hand-built components; end-to-end models collapsed the chain but handled only one task.
- Natively multimodal training on audio, video and text lets the model keep code-switched terms like "mid century" in a Spanish answer with no hand-written rules.
- A three-way tension: raising the thinking budget lifts intelligence scores while time-to-first-audio falls apart.
- Demos include multi-speaker live translation, a roadside-assistance agent, and proactive audio that ignores background noise.
More from Models
- Bug Hunt Bench: multiple runs boost small-model bug detection but move frontier models just 1-2 points — PawelHuryn · 2026-09-15
- ZDTaichu5.0-9B, a 9B vision-language model with spatial reasoning, trends on Hugging Face — TaichuAI · 2026-09-15
- Which 10Eros video-model quant works best on 8GB VRAM? A practical trade-off question — apostrophefee · 2026-09-15
- rasbt shows why final-result benchmarks mislead: Astra vs Qwen in Paint — rasbt · 2026-09-15
- Cristóbal Valenzuela praises Solaris: 'Websites are going to be fun again' — c_valenzuelab · 2026-09-15
- OpenAI books GPT-6 Community Nights in SF and London, confirming model codename Astra — SamuelMLSmith · 2026-09-15