One topological feature boosts noisy speech recognition accuracy from 84.9% to 88.4% on TIMIT

bravo_abad · x · 2026-09-21

Feng et al. show that adding a single manually computed topological feature to a GRU improves noise robustness: speech is converted to a geometric object via time-delay embedding, and persistent homology measures how long the most prominent structure survives across scales. That one number, appended to six learned features for voiced/voiceless consonant classification, raises TIMIT accuracy under strong Gaussian noise from 84.9% to 88.4% while cutting run-to-run variability from ±6.0% to ±0.4% — notable because deep nets are supposed to learn useful features themselves.

Related event: Handcrafted Topological Feature Boosts Speech Network Noise Robustness(2 posts)→

Original post →

More from Research

Research channel →