WaveNet Turns 10: Author Traces Speech Synthesis from Waveforms to Multimodal LLMs
heiga_zen · x · 2026-09-09
Google speech researcher Heiga Zen marks the 10th anniversary of WaveNet, whose raw-waveform generation paradigm replaced concatenative and statistical parametric TTS in 2016.
The decade in review:
- Encoder-decoder models like Tacotron enabled end-to-end text-to-acoustics
- VITS and NaturalSpeech generated waveforms directly without intermediate representations
- AudioLM and VALL-E tokenized speech into LLM frameworks via SoundStream
- Multimodal LLMs like Gemini now seamlessly fuse modalities for natural dialogue
He predicts the next decade: finer emotion/context modeling, ultra-low-latency conversation, and robot integration.
Related event: WaveNet at ten: authors reflect on speech synthesis evolution(2 posts)→
More from Multimodal
- AI-generated virtual influencer stars in a repellent commercial — AIandDesign · 2026-09-09
- SD checkpoint and sampler comparison: JuggernautXL v8 at 26 steps tested across samplers — NickPassig · 2026-09-09
- In the ChatGPT Images 2.5 ad, a woman gets an AI-designed tattoo inked on her arm — yvgh233 · 2026-09-09
- A pixel-art animation site built with Codex and GPT-5, awaiting Gemini 3.0 Pro — Angaisb_ · 2026-09-09
- One prompt turns your selfie into an Arri Alexa editorial shot with GPT Images 2.5 — aziz4ai · 2026-09-09
- Inside fal's H3 Max Director: streaming video generation with mid-stream prompt edits — noahsolomon · 2026-09-09