Alibaba's Qwen3.8-LiveTranslate does real-time speech translation with 2.3s lag and speaker separation
xiaohu · x · 2026-09-19
Alibaba's Qwen team released Qwen3.8-LiveTranslate, a real-time speech-to-speech translation model, with a streaming API (qwen3.8-livetranslate-flash-realtime) on Alibaba Cloud Bailian.
- Thinker–Talker architecture: Thinker interleaves video frames, audio, source text, and translation in one temporal stream; Talker synthesizes translated speech that preserves each speaker's voice, built on a Hybrid MoE backbone.
- Interleave streaming cuts translation lag in the FLEURS benchmark from 2.8s to 2.3s behind the speaker.
- Key features: streaming multi-speaker diarization (spk1/spk2 labels) with per-speaker voice preservation, simultaneous source + translated text output for bilingual subtitles, and disambiguation of names/references using long context and camera frames.
More from Models
- Fruit fly connectome chess model beats Jev 4-1 in 10 games, with a playable demo site — maximelabonne · 2026-09-20
- Bonsai 2 27B safety guardrails reportedly cut SWE-bench and Terminal-bench scores by ~20 points — julianharris · 2026-09-20
- MiMo-V2.6 livestreams its RL run at ~2B tokens/step as Stanford's Marin pretrains in public — stanfordnlp · 2026-09-20
- Jev as a New Primitive: Cheaper Scaling, Better Verifiers and Agent Harness Experiences — omarsar0 · 2026-09-20
- What justifies paying more for an AI model? The 16-second break-even math — zeuslac · 2026-09-20
- Skeptical deep dive confirms Humanity's Last Exam errors; official o3-mini grader marked right answers wrong every time — paul_cal · 2026-09-20