Ten years of WaveNet: Google researcher traces speech synthesis from raw waveforms to multimodal LLMs
heiga_zen · x · 2026-09-09
Google speech researcher Heiga Zen marks the 10th anniversary of WaveNet with a retrospective on a decade of speech synthesis.
- WaveNet (2016): shattered the concatenative/statistical-parametric paradigm by directly generating raw audio waveforms
- End-to-end era: encoder-decoder models like Tacotron generated acoustic features from text, while text-to-waveform models like VITS and NaturalSpeech removed intermediate features entirely
- LLM convergence: AudioLM with SoundStream and VALL-E tokenized audio into the LLM framework, culminating in today's multimodal speech generation
A firsthand technical history from someone who lived it — equal parts tutorial and industry lore.
Related event: WaveNet at ten: authors reflect on speech synthesis evolution(2 posts)→
More from Research
- OpenAI claims internal AI with 10,000 agents solved 90-year-old Navier-Stokes problem — The Verge AI · 2026-09-09
- ValsAI Launches Tax Agent Bench: 391 Expert Questions to Test LLMs on Corporate Tax — JenniferHli · 2026-09-09
- AI Intern reproduces ICML best paper "How much do LLMs memorize" for under $8 — _akhaliq · 2026-09-09
- Buckmaster & Alpöge publish Euler (112pp), Boussinesq (76pp), porous media papers with Lean proofs — xennygrimmato_ · 2026-09-09
- Boltz models keep improving as partners solve previously failed drug-discovery challenges weekly — GabriCorso · 2026-09-09
- Boltz model contributor named to MIT Technology Review's 2026 Innovators Under 35 — GabriCorso · 2026-09-09