Multimodal Attacks Undermine Speaker Anonymization Protections
A new arXiv preprint finds that combining information from multiple anonymized speech utterances—using audio, prosody, and text—significantly improves the ability to verify a speaker's identity, even when anonymization techniques are applied. The study shows that multimodal systems outperform unimodal ones, and that aggregating as few as five anonymized utterances can reduce error rates by over 15% compared to audio-only methods. This suggests that current anonymization methods may not fully protect speaker privacy when adversaries have access to multiple utterances and multimodal data.
Why it matters: The findings highlight a significant privacy risk, indicating that widely used speaker anonymization techniques may be vulnerable to re-identification attacks using multimodal aggregation.
Full story at: arXiv Audio and Speech Processing ↗