First Open-Source Diffusion ASR Model Runs 15x Faster Than Whisper

bodonoghue85 · x · 2026-07-03

interfazeai has open-sourced the first automatic speech recognition (ASR) system based on a diffusion model. By integrating DiffusionGemma and Whisper technologies, it achieves inference speeds 15 times faster than Whisper. This marks the first application of diffusion models in open-source ASR, delivering a significant speed breakthrough while maintaining Whisper's recognition quality, and paving a new technical path for real-time voice applications.

Original post →

More from Models

Models channel →