YODAS v3 lands on Hugging Face: 1.1M hours, the largest open speech dataset ever

shinjiw_at_cmu · x · 2026-09-28

The ESPnet team released YODAS v3 on Hugging Face: 1.1 million hours of audio, the largest open speech dataset ever, and the first at this scale to offer 48kHz stereo audio with timestamped transcripts and translations across 100+ languages, under CC-BY-3.0.

The authors argue training data is the true kingmaker in voice AI: systems like Whisper were built on massive proprietary audio collections, while prior open resources were narrow — LibriLight (60k hours of English) and Multilingual LibriSpeech came from read audiobooks unlike real conversation, and VoxPopuli's multilingual recordings were mostly untranscribed. YODAS v3 aims to close that gap for open voice AI research.

Original post →

More from Research

Research channel →