Alibaba's Qwen3.8-LiveTranslate cuts speech translation lag to 2.3s, adds speaker separation

xiaohu · x · 2026-09-19

Alibaba released Qwen3.8-LiveTranslate, a real-time speech translation model that listens, translates, and speaks simultaneously. Key upgrades: streaming speaker separation (spk1/spk2/spk3 labels with per-speaker voice preservation), simultaneous source-transcript and translation output for subtitles, disambiguation using long context plus camera frames for homophones and named entities, and reduced streaming latency — average translation lag on FLEURS dropped from 2.8s to 2.3s. Coverage: 60 languages for audio input, 60 for text output, 29 for speech output.

Related event: Alibaba's Qwen releases LiveTranslate with 2.3-second latency across 60 languages(4 posts)→

Original post →

More from Models

Models channel →