YODAS v3 lands on Hugging Face: 1.1M hours, the largest open speech dataset ever
shinjiw_at_cmu · x · 2026-09-28
The ESPnet team released YODAS v3 on Hugging Face: 1.1 million hours of audio, the largest open speech dataset ever, and the first at this scale to offer 48kHz stereo audio with timestamped transcripts and translations across 100+ languages, under CC-BY-3.0.
The authors argue training data is the true kingmaker in voice AI: systems like Whisper were built on massive proprietary audio collections, while prior open resources were narrow — LibriLight (60k hours of English) and Multilingual LibriSpeech came from read audiobooks unlike real conversation, and VoxPopuli's multilingual recordings were mostly untranscribed. YODAS v3 aims to close that gap for open voice AI research.
More from Research
- Sakana AI founder hardmaru's Neuroevolution textbook is finally in print — hardmaru · 2026-09-28
- Free video course list: build LLMs, DeepSeek and diffusion models from scratch — caglar_ee · 2026-09-28
- Hillock: open-source neuro-symbolic agent memory engine runs under 1.2GB VRAM — Equivalent-Flan-1590 · 2026-09-28
- PKU open-sources RayOrch, lineage-aware data-prep engine with up to 15.14x speedup — PekingUniversity · 2026-09-28
- Graph alignment is all you need: slides from CIRM workshop talk — marc_lelarge · 2026-09-28
- AI-picked catalyst dismissed by experts survived 1,000+ hours in acid — VraserX · 2026-09-28