RobustSpeechFlow: Contrastive Flow Matching Reduces TTS Errors
RobustSpeechFlow is a new training strategy for text-to-speech (TTS) systems that extends contrastive flow matching with length-preserving augmentations to simulate repeat and skip errors. This approach improves alignment robustness and reduces content fidelity errors without requiring external aligners or preference data. On the Seed-TTS-eval benchmark, RobustSpeechFlow reduces word error rate from 1.44 to 1.38 using only 0.06B parameters. On the ZERO500 benchmark, it lowers English character error rate from 0.48% to 0.35% and Korean from 0.81% to 0.57%.
Why it matters: This method offers a practical way to improve TTS content fidelity and robustness, making it easier to integrate into existing pipelines without extra data or tools.
Full story at: arXiv Audio and Speech Processing ↗