otoSpeech Task releases 20 hours of full-duplex task-oriented conversation data
kastnerkyle · x · 2026-09-02
oto.earth has released the otoSpeech Task dataset, featuring 20 hours of English two-person task-oriented conversations under the CC BY 4.0 license.
Dataset Features
- Full-Duplex & Context: Retains the images speakers saw, actions taken, and reference answers, all aligned on a single timeline.
- Rich Metadata: Includes 58 sessions across seven collaborative tasks with 11,025 timestamped events.
Context
The post contrasts cascaded systems (ASR+LLM+TTS) with full-duplex models. Cascaded systems offer inspectable task states but lack natural flow, while full-duplex models feel natural but lag in task reliability. This dataset aims to bridge the gap, advancing voice AI toward systems that talk like humans while maintaining task accuracy.
More from Research
- MirroS Introduces Code-as-World: Representing the Physical World as Executable Code — le_james94 · 2026-09-02
- OpenAI's Astra achieves 100% success rate on ExploitBench vulnerability tests — JiaweiLiu_ · 2026-09-02
- MLPerf Storage v3.0 lands with 144 results, adding KV cache and vector DB tests — TheKanter · 2026-09-02
- Frontier Labs Update: ExploitBench on Open-Source Benchmarks — moyix · 2026-09-02
- Anthropic Reports Incidents of Models Gaining Unauthorized Access — rickasaurus · 2026-09-02
- Developer finds flaws in CritPt benchmark scores — scaling01 · 2026-09-02