One audio encoder + LLM serving offline and streaming ASR without accuracy tax
rohanpaul_ai · x · 2026-09-17
Rohan Paul highlights the architecture of NetEase Youdao's Confucius4-R2T2: a single Audio Encoder + LLM foundation serves both offline and streaming recognition, with configurable chunk sizes and no accuracy tax offline — one system tuned along a latency/quality axis instead of two drifting models. He argues generic ASR collapses exactly where real work happens: a meeting assistant that can't spell your internal tool names quietly becomes unusable in the setting it was built for.
Related event: NetEase Youdao Open-Sources Confucius4-R2T2 Streaming ASR Model(7 posts)→
More from Models
- Databricks rolls out GPT-6 Astra to 3,500 engineers, coding spend jumps 60% — nickbaumann_ · 2026-09-17
- How do models know if they're in training, evaluation, or real inference? Nobody really knows — flowersslop · 2026-09-17
- Google brings Gemma 4 12B to Mac, running locally on 16GB MacBook Air — GlennCameronjr · 2026-09-17
- DeepSeek ends Google's 51-week OpenRouter streak with 12.6× request growth to 1.01B weekly — rohanpaul_ai · 2026-09-17
- Qwen 3.8's structurally correct but nonsensical replies baffle Reddit users — TastesLikeOwlbear · 2026-09-17
- Power user says Opus understands their intent in a way Fable never quite does — RileyRalmuto · 2026-09-17