415k hours of full-duplex dialogue speech dataset released for spoken dialogue model training

kastnerkyle · x · 2026-09-10

wataru9871 released 415k hours of full-duplex dialogue speech data, aimed at training full-duplex spoken dialogue models, with an accompanying paper and open-source code. It's one of the largest public resources of its kind for real-time voice interaction research.

Original post →

More from Research

Research channel →