Do Speech Tokens Leak Voiceprints? Speaker Inversion Attacks Against End-to-End Speech Language Models
A new preprint introduces Audio BERT (AuB) and SpInv, a two-stage method for recovering speaker embeddings from speech tokens produced by end-to-end speech language models. The study demonstrates that SpInv can achieve cosine similarities above 0.70 with just three seconds of speech token output from models such as Moshi and Qwen3-Omni, indicating that these tokens can leak significant speaker identity information.
Why it matters: This work highlights a substantial privacy risk in modern speech language models, showing that speech tokens can be inverted to recover speaker voiceprints and potentially compromise user anonymity.
Full story at: arXiv Cryptography and Security ↗