Alibaba open-sources two speech enhancement models: 1.5-2.3x faster separation, 22ms echo cancellation
aigclink · x · 2026-09-10
Alibaba open-sourced two new AI speech enhancement models, completing a suite of seven covering background noise, echo, and overlapping speakers. FLASepformer handles speech separation with linear attention, cutting quadratic compute to linear—1.5-2.3x faster with only 16-32% of the original memory. JAEC 16K is an interpretable echo canceller whose internal steps (delay estimation, echo estimation and suppression) are observable for debugging, with just 22ms latency for real-time calls; it doesn't yet handle nonlinear distortion like speaker clipping. Weights for both are released.
More from Models
- Pedro Domingos: It's Time to Call Residual Streams What They Are — RAM — pmddomingos · 2026-09-10
- Japanese doctor recreates coronary angiography teaching demo with GPT-6 Astra — TheMoonMidas · 2026-09-10
- Iron Man fan builds interactive suit teardown site with GPT-6 Astra, Three.js — TheMoonMidas · 2026-09-10
- Google launches Gemini 3.8 Flash Cyber, a frontier cybersecurity model that finds and patches vulnerabilities autonomously — OfficialLoganK · 2026-09-10
- Why can't I pay OpenAI or Anthropic to hack my systems with unsafe models, asks David Holz — DavidSHolz · 2026-09-10
- Pixel analysis of OpenAI's Navier-Stokes chart suggests ~91 solved open problems kept unreleased — teortaxesTex · 2026-09-10