NetEase Youdao open-sources Confucius R2T2: 160ms streaming slices for speech AI

VraserX · x · 2026-09-17

NetEase Youdao has open-sourced Confucius R2T2, and its headline number of 160ms is not end-to-end latency but the minimum streaming audio step — an important distinction. The real highlight is that the model processes speech continuously in tiny slices while you're still talking, rather than waiting for a big chunk of audio. This always-on streaming understanding is the kind of change that makes voice AI feel like an actual conversation.

Related event: NetEase Youdao Open-Sources Confucius R2T2 Speech Model(2 posts)→

Original post →

More from Multimodal

Multimodal channel →