415k hours of full-duplex dialogue speech dataset released for spoken dialogue model training
kastnerkyle · x · 2026-09-10
wataru9871 released 415k hours of full-duplex dialogue speech data, aimed at training full-duplex spoken dialogue models, with an accompanying paper and open-source code. It's one of the largest public resources of its kind for real-time voice interaction research.
More from Research
- The Prism Hypothesis unified autoencoding paper accepted at ECCV 2026, paves way for encoder-free MLLMs — liuziwei7 · 2026-09-10
- embedflow Migrates Embedding Models Without Re-embedding: Qwen 4B→8B Matches Native Retrieval with Just 50 Reranked Docs — Potential_Low_1183 · 2026-09-10
- Degenerate Fisher information explains why huge neural nets don't defy Occam's razor — FrnkNlsn · 2026-09-10
- AIxBio researcher: skip the bitter lesson debate — more compute means lower per-dollar efficiency — anshulkundaje · 2026-09-10
- GPN-Star genomics paper hailed as the blueprint for good AIxBio research — anshulkundaje · 2026-09-10
- Browser demo of TAMER makes human-feedback RL policy changes tangible — PeterStone_TX · 2026-09-10