Test: Room Reverberation and Low SNR Hurt STT More Than Model Size

ChromaForge · reddit · 2026-07-30

In Speech-to-Text (STT) tasks, when transcription fails, people often blame the model size while ignoring the physical audio input quality.

Through acoustic testing, the author found that room reverberation (comb filtering) and low-frequency environmental noise (lifting the noise floor) severely degrade high-frequency energy, causing token decoding errors. Experiments show that applying adaptive spectral subtraction and front-end DSP filtering to improve the Signal-to-Noise Ratio (SNR) prior to decoding yields a significantly greater Word Error Rate (WER) reduction than simply upgrading from Whisper-medium to Large-v3.

Original post →

More from Research

Research channel →