WaveNet Turns 10: Author Traces Speech Synthesis from Waveforms to Multimodal LLMs

heiga_zen · x · 2026-09-09

Google speech researcher Heiga Zen marks the 10th anniversary of WaveNet, whose raw-waveform generation paradigm replaced concatenative and statistical parametric TTS in 2016.

The decade in review:

He predicts the next decade: finer emotion/context modeling, ultra-low-latency conversation, and robot integration.

Related event: WaveNet at ten: authors reflect on speech synthesis evolution(2 posts)→

Original post →

More from Multimodal

Multimodal channel →