NetEase Youdao open-sources Confucius4-R2T2 streaming ASR with 200-600ms latency
aigclink · x · 2026-09-20
NetEase Youdao has released Confucius4-R2T2, an open-source real-time speech recognition model built on Qwen3-ASR, designed to eliminate text flickering in live captioning.
- Append-only streaming output: once committed, transcript text is never revised, avoiding disruptive rewrites in real-time subtitles.
- LSP strategy: the model keeps listening and only finalizes a segment when it's confident it won't change; uncertain parts wait for more context, balancing stability against latency (200-600ms).
- Configurable decoding chunks from 80ms to 2s: smaller chunks for lower latency, larger chunks for higher accuracy.
Target use cases include live captioning, simultaneous interpretation, voice customer service, and downstream NLP/LLM agent pipelines. Code and WebSocket server/client examples are on GitHub (netease-youdao/Confucius4-R2T2).
Related event: NetEase Youdao Open-Sources Confucius4-R2T2 Real-Time ASR(2 posts)→
More from Models
- StepFun Launches Step 5 Preview: 600B MoE with 27B Active, Open Weights on Oct 15 — airesearch12 · 2026-09-20
- LLM vs classical ML across 8 datasets: labeled data still favors SVM and XGBoost — Ok_Juggernaut2187 · 2026-09-20
- Researchers Use Claude Opus 5 to Hijack OpenAI Employee Accounts for Under $3,000 — FinanceYF5 · 2026-09-20
- Alibaba open-sources DAMO RADAR medical model detecting cancer and ~150 conditions — emmanuelvivier · 2026-09-20
- Anthropic reportedly testing Claude Money to link bank accounts for spending analysis — emmanuelvivier · 2026-09-20
- Mozilla report: open-weight AI models now just four months behind the closed frontier — mark_k · 2026-09-20