Multimodal Speaker Verification as a Threat to Speaker Anonymization
TL;DR AI
2 min readKey summary
A new study finds anonymized speech can still be used to verify speakers when systems combine multiple utterances with audio, prosodic, and linguistic cues.
Accuracy improves as more anonymized utterances are aggregated, showing speaker identity leaks persist beyond single-clip analysis.
Multimodal and frame-level aggregation methods outperform audio-only approaches, lowering equal error rate and strengthening verification.
The results suggest anonymization methods focused only on vocal traits may not fully protect privacy in speech systems.
