AirCaps Launches Audio Research Lab, Citing 20% WER for ASR on Noisy Real-World Speech
ycombinator · x · 2026-09-03
Voice startup AirCaps publicly launched the AirCaps Audio Research Lab, targeting messy real-world acoustics: far-field speech, reverberation, low SNR, and overlapping multi-speaker audio.
Key facts:
- Cites Ayllon et al. (2026): 40 leading ASR models hit a median 20.45% word error rate on noisy speech, versus the 2-3% they report on clean benchmarks.
- Largest speech datasets and benchmarks (e.g., Librispeech) fall far short of real-world complexity; real noisy audio data remains scarce.
- AirCaps builds fully streaming speech/audio models — target speaker extraction, far-field enhancement, separation, diarization, speech-to-text — claiming a step-function lead over major cloud speech models while running on-device on consumer hardware.
The team has 8 years of consumer-electronics voice R&D experience.
More from Models
- Baseten ships GLM-5.3 Fast: speed-optimized open-weight model for real-time workloads — baseten · 2026-09-03
- Users Report Claude Racking Up Daily Mistakes and Hallucinations — lilyraynyc · 2026-09-03
- First run of Gemini 3.8 Flash fails: model keeps thinking until it times out — rickasaurus · 2026-09-03
- Early user verdict: fable 5 outperforms the newer fable 5.1 — BLUECOW009 · 2026-09-03
- Flash 3.8 Review: Great When Working, but Stuck in Silent Token-Burning Loops — brandon_galang · 2026-09-03
- Claude Max 20x buyer says weekly limits, not the 5-hour window, are the real bottleneck; r/ClaudeAI deleted his post — conorearly · 2026-09-03