WhisperVC: Decoupled Cross-Domain Alignment and Speech Generation for Low-Resource Whisper-to-Normal Conversion
WhisperVC is a three-stage framework designed to convert whispered speech to normal speech using limited paired data. By separating cross-domain alignment from speech generation, the method achieves competitive results, with DNSMOS 3.07 and WavLM speaker similarity of 0.95 on the AISHELL6-Whisper dataset. The approach demonstrates potential for privacy-preserving communication and as an assistive tool for individuals with vocal-fold impairments.
Why it matters: This work advances whisper-to-normal speech conversion in low-resource settings, enabling practical applications in privacy and assistive technology.
Full story at: arXiv Audio and Speech Processing ↗