Alibaba open-sources 2 more speech enhancement models, completing a 7-model toolkit with 22ms echo cancellation

aigclink · x · 2026-09-10

Alibaba open-sourced two new AI speech enhancement models, completing a 7-model system covering noise, echo, and overlapping speech: FLASepformer uses linear attention to turn quadratic-cost speech separation into linear (1.5-2.3x faster, 16-32% of the original memory), while JAEC 16K makes echo cancellation observable — its two internal steps (latency estimation, echo estimation/cancellation) are inspectable — with just 22ms latency, though it only handles linear echo, not non-linear distortion. The post also maps all 7 models (ZipEnhancer, DFSMN ANS, FRCRN, MossFormer2, DFSMN AEC) to scenarios: noise-reduction stacks for assistants, JAEC+denoise for real-time calls, separation+VAD/ASR for multi-speaker meetings.

Related event: Alibaba Open-Sources Two Speech Enhancement Models, Completing 7-Model Suite(2 posts)→

Original post →

More from Multimodal

Multimodal channel →