Xiaomi open-sources CocktailASR-1, a target-speaker ASR hitting SOTA on multi-speaker benchmarks
solyarisoftware · x · 2026-09-17
Xiaomi released and open-sourced Xiaomi-CocktailASR-1, an end-to-end target-speaker ASR model that locks onto one reference voiceprint and transcribes only that speaker in crowded, overlapping audio.
- Multi-speaker SOTA: WER 4.11% on LibriMix 2mix, crushing Qwen3-ASR (68.75%) and Gemini (48.41%)
- Single-speaker parity: same model handles both scenarios, no ASR switch needed
- Negative-sample rejection: near-zero false triggers on non-target speech
- Chain-of-Thought reasoning for explainable transcription
- Supports English and Chinese. Repo: xiaomi-research/xiaomi-cocktailasr-1
More from Models
- LAB legal benchmark flaw: case docs leak planted issues directly to models — andersonbcdefg · 2026-09-17
- MiMo eval chart shows judge and probe disagree 60% of the time, sparking reward-hacking concerns — andrew_n_carr · 2026-09-17
- Fans mourn the end of Codex lead's regular daily reset cadence — kimmonismus · 2026-09-17
- Users report Google Astra feels noticeably degraded over past two days — pwlot · 2026-09-17
- OpenAI burns 20% as much compute on monitoring as the model itself, SemiAnalysis says — kevinnbass · 2026-09-17
- Google releases Gemma 3n: 2GB RAM multimodal model, first sub-10B to top 1300 on LMArena — joemeno · 2026-09-17