Switch language한국어
Back to the list

Voice Memory for Agentic Speech Recognition

TL;DR AI

Key summary

2 min read
  1. Researchers introduced Voice Memory, an inference-only listener-thinker architecture for agentic speech recognition.

  2. A frozen corrector consults a per-domain memory file at inference time, while a score-gated optimizer updates memory only when edits improve a held-out metric.

  3. The approach reduced overcorrection versus unconstrained generative error correction and lowered weighted word error rate across ten HyPoradise domains.

  4. It delivered especially strong gains on air-travel commands and noisy far-field speech, with no added inference-time parameters.

Read the original