One audio encoder + LLM serving offline and streaming ASR without accuracy tax

rohanpaul_ai · x · 2026-09-17

Rohan Paul highlights the architecture of NetEase Youdao's Confucius4-R2T2: a single Audio Encoder + LLM foundation serves both offline and streaming recognition, with configurable chunk sizes and no accuracy tax offline — one system tuned along a latency/quality axis instead of two drifting models. He argues generic ASR collapses exactly where real work happens: a meeting assistant that can't spell your internal tool names quietly becomes unusable in the setting it was built for.

Related event: NetEase Youdao Open-Sources Confucius4-R2T2 Streaming ASR Model(7 posts)→

Original post →

More from Models

Models channel →