Alibaba open-sources two speech enhancement models: 1.5-2.3x faster separation, 22ms echo cancellation

aigclink · x · 2026-09-10

Alibaba open-sourced two new AI speech enhancement models, completing a suite of seven covering background noise, echo, and overlapping speakers. FLASepformer handles speech separation with linear attention, cutting quadratic compute to linear—1.5-2.3x faster with only 16-32% of the original memory. JAEC 16K is an interpretable echo canceller whose internal steps (delay estimation, echo estimation and suppression) are observable for debugging, with just 22ms latency for real-time calls; it doesn't yet handle nonlinear distortion like speaker clipping. Weights for both are released.

Related event: Alibaba Open-Sources Two Speech Enhancement Models, Completing 7-Model Suite(2 posts)→

Original post →

More from Models

Models channel →