NetEase Youdao open-sources Confucius R2T2: 160ms streaming slices for speech AI
VraserX · x · 2026-09-17
NetEase Youdao has open-sourced Confucius R2T2, and its headline number of 160ms is not end-to-end latency but the minimum streaming audio step — an important distinction. The real highlight is that the model processes speech continuously in tiny slices while you're still talking, rather than waiting for a big chunk of audio. This always-on streaming understanding is the kind of change that makes voice AI feel like an actual conversation.
Related event: NetEase Youdao Open-Sources Confucius R2T2 Speech Model(2 posts)→
More from Multimodal
- Grok Imagine adds in-image text editing, no more regenerating posters over one wrong word — XFreeze · 2026-09-17
- YuE2 open-source music model goes head-to-head with Suno and ElevenLabs in blind comparison — cocktailpeanut · 2026-09-17
- Tutorial: Turn MiniMax AI Video Into Gaussian Splatting 3D via COLMAP — Many-Ad-6225 · 2026-09-17
- Higgsfield launches API with 50+ models, open-sources $5.4B startup's code, offers $50k bounty — SimplyAnnisa · 2026-09-17
- Grok Voice goes live on fal: 0.70s latency, word-level timestamps, 2-minute voice cloning — SpaceXAI · 2026-09-17
- Image model Jev redraws MaskGIT from first principles by predicting every pixel in parallel — torchcompiled · 2026-09-17