Voice Memory for Agentic Speech Recognition
TL;DR AI
2 min readKey summary
Researchers introduced Voice Memory, an inference-only listener-thinker architecture for agentic speech recognition.
A frozen corrector consults a per-domain memory file at inference time, while a score-gated optimizer updates memory only when edits improve a held-out metric.
The approach reduced overcorrection versus unconstrained generative error correction and lowered weighted word error rate across ten HyPoradise domains.
It delivered especially strong gains on air-travel commands and noisy far-field speech, with no added inference-time parameters.
