WaveNet turns 10: audio-as-language-modeling is now the standard, dilated convs still everywhere
heiga_zen · x · 2026-09-09
A researcher reshared DeepMind's original WaveNet announcement to mark the 10th anniversary of the landmark paper.
- Audio generation framed as language modeling has become the standard paradigm, and the paper's dilated convolutions still appear throughout modern codecs and vocoders.
- The author recalls being struck a decade ago that this would transform speech synthesis forever, while noting the paper was written very 'frugally' — details and ablations felt scarce because, per one author, the team rushed the writeup to share it as widely and quickly as possible.
Related event: WaveNet Marks 10th Anniversary as Foundational Audio Generation Paradigm(4 posts)→
More from Research
- The Mathematics Autoformalization Project: translating all known math into formal code — burny_tech · 2026-09-09
- OpenFrontier (RSS 2026): zero-shot open-world robot navigation with VLM-scored frontiers — rsasaki0109 · 2026-09-09
- Open problems turned into RL environments: benchmarks and RL envs are two sides of the same coin — burny_tech · 2026-09-09
- ICML 2026 outstanding paper drama: concurrent diffusion sampling result, only one got the award — peter_richtarik · 2026-09-09
- UrbanLLMind: 10k memory-equipped LLM agents simulate a week of real San Francisco movement — anas_ant · 2026-09-09
- Radial Science commits $20M to Prism to make protein motion measurable and actionable — anshulkundaje · 2026-09-09