Alibaba's Qwen3.8-LiveTranslate cuts speech translation lag to 2.3s, adds speaker separation
xiaohu · x · 2026-09-19
Alibaba released Qwen3.8-LiveTranslate, a real-time speech translation model that listens, translates, and speaks simultaneously. Key upgrades: streaming speaker separation (spk1/spk2/spk3 labels with per-speaker voice preservation), simultaneous source-transcript and translation output for subtitles, disambiguation using long context plus camera frames for homophones and named entities, and reduced streaming latency — average translation lag on FLEURS dropped from 2.8s to 2.3s. Coverage: 60 languages for audio input, 60 for text output, 29 for speech output.
More from Models
- Fruit fly connectome chess model beats Jev 4-1 in 10 games, with a playable demo site — maximelabonne · 2026-09-20
- Bonsai 2 27B safety guardrails reportedly cut SWE-bench and Terminal-bench scores by ~20 points — julianharris · 2026-09-20
- MiMo-V2.6 livestreams its RL run at ~2B tokens/step as Stanford's Marin pretrains in public — stanfordnlp · 2026-09-20
- Jev as a New Primitive: Cheaper Scaling, Better Verifiers and Agent Harness Experiences — omarsar0 · 2026-09-20
- What justifies paying more for an AI model? The 16-second break-even math — zeuslac · 2026-09-20
- Skeptical deep dive confirms Humanity's Last Exam errors; official o3-mini grader marked right answers wrong every time — paul_cal · 2026-09-20