NVIDIA's Voice Memory: Optimizing Speech Recognition via Editable Memory Files

nvidia · hf · 2026-07-31

NVIDIA introduced Voice Memory, an inference-only scheme for agentic speech recognition utilizing a listener-thinker architecture:

This design requires no weight changes, keeping learned skills auditable and portable. Experiments show that unconstrained generative error correction often over-corrects, breaking correct tokens (up to 64% of the time on financial news), whereas Voice Memory reduces this rate to 35%. Across ten domains, the scheme lowers the weighted word error rate (WWER) from 8.36% to 7.52% without adding inference parameters.

Original post →

More from Research

Research channel →