Alibaba open-sources 2 more speech enhancement models, completing a 7-model toolkit with 22ms echo cancellation
aigclink · x · 2026-09-10
Alibaba open-sourced two new AI speech enhancement models, completing a 7-model system covering noise, echo, and overlapping speech: FLASepformer uses linear attention to turn quadratic-cost speech separation into linear (1.5-2.3x faster, 16-32% of the original memory), while JAEC 16K makes echo cancellation observable — its two internal steps (latency estimation, echo estimation/cancellation) are inspectable — with just 22ms latency, though it only handles linear echo, not non-linear distortion. The post also maps all 7 models (ZipEnhancer, DFSMN ANS, FRCRN, MossFormer2, DFSMN AEC) to scenarios: noise-reduction stacks for assistants, JAEC+denoise for real-time calls, separation+VAD/ASR for multi-speaker meetings.
More from Multimodal
- 3D Gaussian splat reconstruction of Amazon Prime Air crash site from NTSB footage — bilawalsidhu · 2026-09-10
- 3D Gaussian Splat Rebuilds Amazon Prime Air Crash Site from New NTSB Footage — bilawalsidhu · 2026-09-10
- OpenAI case study: GPT-6 Astra builds a house in Blender from one prompt, then moves it into UE5 — xiaohu · 2026-09-10
- OpenAI demos Astra driving Blender to build editable 3D scenes from prompts — xiaohu · 2026-09-10
- One prompt to an interactive UE5 home: OpenAI's GPT-6 Astra 3D design case — xiaohu · 2026-09-10
- VivagoR1 launches as a conversational AI video agent promising deterministic 5-minute outputs — 量子位 · 2026-09-10