Alibaba's Qwen Audio 3.0 TTS Plus has topped the Speech Arena leaderboard by Artificial Analysis. The model supports 16 languages and allows users to control speaking style via natural language or tags like [angry], but it is significantly slower than rivals, generating only 16 characters per second.
Why it matters: This model sets a new quality benchmark in text-to-speech, but its slow speed highlights the trade-off between quality and latency in AI voice generation.
Alibaba's Tongyi Lab has released Qwen-Audio-3.0-TTS, a production-oriented text-to-speech system available in Flash (real-time) and Plus (high-quality) tiers. The model supports 16 languages and is delivered as a hosted service via Alibaba Cloud Model Studio, rather than as downloadable weights.
Why it matters: This release offers developers a scalable, hosted TTS solution optimized for production use cases across multiple languages.
Research→Official→arXiv Audio and Speech Processing
FlowSonic is a zero-shot music editing framework that leverages a pretrained diffusion transformer with rectified flow to enable text-guided editing of real-world music recordings. The method uses deterministic inversion and cross-attention reuse to preserve musical structure, while a high-order ODE solver enhances numerical stability during editing. Experiments indicate that FlowSonic outperforms existing approaches in tasks such as timbre transfer and genre modification, with improvements in semantic alignment, structural consistency, and perceptual audio quality.
Why it matters: This work demonstrates a significant advance in controllable music editing without task-specific training, potentially impacting music production and audio post-processing workflows.
Research→Official→arXiv Audio and Speech Processing
WhisperVC is a three-stage framework designed to convert whispered speech to normal speech using limited paired data. By separating cross-domain alignment from speech generation, the method achieves competitive results, with DNSMOS 3.07 and WavLM speaker similarity of 0.95 on the AISHELL6-Whisper dataset. The approach demonstrates potential for privacy-preserving communication and as an assistive tool for individuals with vocal-fold impairments.
Why it matters: This work advances whisper-to-normal speech conversion in low-resource settings, enabling practical applications in privacy and assistive technology.
Research→Official→arXiv Audio and Speech Processing
UniPASE is an extension of the PASE framework designed for universal speech enhancement across multiple sampling rates. It introduces the DeWavLM-Omni module, fine-tuned from WavLM via knowledge distillation, to convert degraded waveforms into clean, linguistically faithful phonetic representations. The system reconstructs high-fidelity 16 kHz waveforms and upscales them to 48 kHz, achieving robust enhancement with minimal linguistic hallucination. UniPASE demonstrated superior or competitive performance compared to state-of-the-art models and won first place in the objective evaluation of the URGENT 2026 Challenge.
Why it matters: UniPASE advances the field of speech enhancement by addressing both fidelity and hallucination issues across diverse sampling rates, setting a new benchmark in universal speech enhancement.
Research→Official→arXiv Audio and Speech Processing
RealDESED is a new benchmark dataset for domestic sound event detection, featuring 5,710 real-world audio recordings from 652 participants' homes. The dataset includes 15 sound classes, multi-annotator labels, and extensive metadata, aiming to better reflect the variability and complexity of real domestic environments compared to simulated datasets. A transformer-based baseline achieves a macro-averaged PSDS1 score of 0.731 on the test set.
Why it matters: RealDESED offers a realistic and diverse dataset that can drive the development of more robust sound event detection systems for real-world home environments.
Research→Official→arXiv Audio and Speech Processing
SALMONN-2 is an audio large language model (ALLM) built on a unified self-supervised learning (SSL) encoder, enhanced by a multi-layer feature fusion adapter that aggregates hierarchical representations. The model achieves state-of-the-art performance on several ALLM understanding benchmarks among comparable-scale open-weight models. The study also demonstrates that multimodal in-context learning (MICL) can be effectively acquired through targeted contextual biasing training, rather than emerging naturally.
Why it matters: This work demonstrates that general-purpose self-supervised audio encoders can match or surpass specialized supervised encoders, potentially streamlining the development of versatile audio AI systems.
Research→Official→arXiv Audio and Speech Processing
Researchers introduce NABEATs, a noise-aware audio self-supervised learning model that leverages a reference noise input to help suppress undesired noise in audio representations. NABEATs is trained to estimate clean audio representations from noisy signals, leading to improved performance on downstream tasks in noisy environments and better generalization to previously unseen noise types.
Why it matters: This work advances the robustness of audio AI systems in real-world noisy conditions, which is important for applications such as speech recognition and sound event detection.
Research→Official→arXiv Audio and Speech Processing
ChipChat introduces a novel low-latency cascaded system for real-time on-device voice agents, integrating streaming speech recognition, large language models, text-to-speech, vocoder, and speaker modeling. Implemented in MLX, the system achieves sub-second response latency on a Mac Studio without dedicated GPUs, enabling privacy-preserving, fully on-device processing. The work demonstrates that architectural innovations and streaming optimizations can overcome traditional latency bottlenecks in cascaded systems.
Why it matters: This research provides a practical solution for real-time, privacy-preserving voice-based AI agents on consumer hardware, addressing a key challenge in deploying conversational AI locally.
Research→Official→arXiv Audio and Speech Processing
A preprint introduces Precision-Varying Prediction (PVP), a method that improves the adversarial robustness of automatic speech recognition (ASR) models by randomly varying the inference precision. The approach also enables adversarial example detection using a Gaussian classifier on outputs from different precision levels. When combined with uncertainty-based defenses, PVP increases the difficulty for adaptive attackers, requiring them to introduce more perceptible noise to evade detection. Experimental results show significant improvements in robustness and detection across multiple ASR models, languages, and attack types.
Why it matters: This work proposes a practical and effective defense against adversarial attacks on ASR systems, which are increasingly used in real-world automated applications.
Research→Official→arXiv Audio and Speech Processing
Researchers present X-Translator, a modular, low-cost, open-source speech-to-speech translation system that integrates streaming automatic speech recognition (ASR), machine translation, and prompt-conditioned text-to-speech (TTS). The system is designed for real-time translation in long-form and multi-speaker conversations, addressing challenges such as unstable ASR hypotheses, ambiguous turn boundaries, and speaker consistency. X-Translator is evaluated on translation quality, speech quality, latency, and speaker preservation, with code and a demo publicly available.
Why it matters: X-Translator offers an open, reproducible platform for real-time multilingual speech translation with speaker awareness, helping advance practical deployment in complex conversational scenarios.
Research→Official→arXiv Audio and Speech Processing
Researchers introduce WildElder, a Mandarin elderly speech corpus collected from online videos and annotated with transcription, speaker age, gender, and accent strength. The dataset addresses the scarcity of diverse, real-world elderly speech data for automatic speech recognition and speaker profiling. Experimental results demonstrate the challenges of elderly speech recognition and establish WildElder as a new benchmark for the field.
Why it matters: WildElder provides a much-needed resource for developing and evaluating speech technologies tailored to aging populations.
Policy & Safety→Official→arXiv Cryptography and Security
A new preprint introduces Audio BERT (AuB) and SpInv, a two-stage method for recovering speaker embeddings from speech tokens produced by end-to-end speech language models. The study demonstrates that SpInv can achieve cosine similarities above 0.70 with just three seconds of speech token output from models such as Moshi and Qwen3-Omni, indicating that these tokens can leak significant speaker identity information.
Why it matters: This work highlights a substantial privacy risk in modern speech language models, showing that speech tokens can be inverted to recover speaker voiceprints and potentially compromise user anonymity.
Researchers introduce ESCUCHA, the first Spanish speech understanding benchmark designed to evaluate large audio language models across diverse acoustic conditions and reasoning abilities. The benchmark features 1,000 human-curated questions paired with 162.9 hours of real-world audio, encompassing multiple Spanish accents and non-normative speech. Initial evaluations show substantial performance gaps between state-of-the-art models and trained humans.
Why it matters: This benchmark fills a critical gap in evaluating audio language models for Spanish under realistic and challenging conditions, highlighting the need for more robust and capable multilingual speech AI systems.
Researchers have introduced JarvisBench, a new benchmark designed to evaluate always-on spoken mediators that facilitate real-time interaction between users and long-horizon AI agents. JarvisBench features two tracks: one assessing whether mediation improves agent task completion, and another measuring user understanding and responsiveness. Initial experiments using a modular Jarvis prototype on 34 WildClaw tasks indicate that spoken mediation can enhance both task performance and user comprehension, though its effectiveness is highly dependent on the underlying language model.
Why it matters: JarvisBench provides a standardized framework to assess the impact of real-time spoken mediation in AI agent workflows, addressing a key gap in user-agent interaction.
Sony Music Entertainment has filed a lawsuit against AI music generator Udio, alleging copyright infringement of more than 30,000 songs, including works by Elvis Presley, Beyoncé, and Harry Styles. The lawsuit was filed in a New York court on Monday.
Why it matters: This case could set a major precedent for how copyright law applies to AI-generated music and the liability of AI companies for training on copyrighted works.
Research→Official→arXiv Audio and Speech Processing
A study of 1,168 Japanese voice actors demonstrates that AI voice-clone attribution systems based on embedding similarity have a persistent misidentification floor of approximately 2.6%, even when using advanced ensemble methods. Under session-disjoint conditions, this error rate rises to 13%. Generic encoders cause about half of clones from non-enrolled individuals to be wrongly attributed to enrolled actors, while 32% of clones of enrolled targets are missed. Domain-matched encoders reduce, but do not eliminate, these attribution errors, revealing that fixed-threshold attribution remains unreliable and potentially unfair.
Why it matters: This research highlights a fundamental limitation in current voice-clone attribution systems, raising concerns about the protection and fair treatment of professional voice actors.
Research→Official→arXiv Audio and Speech Processing
RobustSpeechFlow is a new training strategy for text-to-speech (TTS) systems that extends contrastive flow matching with length-preserving augmentations to simulate repeat and skip errors. This approach improves alignment robustness and reduces content fidelity errors without requiring external aligners or preference data. On the Seed-TTS-eval benchmark, RobustSpeechFlow reduces word error rate from 1.44 to 1.38 using only 0.06B parameters. On the ZERO500 benchmark, it lowers English character error rate from 0.48% to 0.35% and Korean from 0.81% to 0.57%.
Why it matters: This method offers a practical way to improve TTS content fidelity and robustness, making it easier to integrate into existing pipelines without extra data or tools.
Policy & Safety→Reported→arXiv Cryptography and Security
Researchers have shown that speech recognition models can be compromised using ordinary natural sounds as backdoor triggers. Their experiments demonstrate that with only 5% of training data poisoned, these attacks achieve nearly 100% success while leaving model performance on benign inputs unaffected. The use of common sounds makes the attacks stealthy and difficult to detect.
Why it matters: This finding exposes a significant security vulnerability in speech recognition systems, as undetectable backdoors could be triggered by everyday sounds.
Research→Official→arXiv Audio and Speech Processing
Researchers present Dialogs, a 20.6-hour Russian conversational speech corpus recorded in a professional studio, featuring 3 speakers and 11,796 utterances. The dataset is notable for capturing turn-taking rhythm and expressive prosody, with per-utterance style and emotion labels across 12 categories. Crowd MOS tests indicate that Dialogs achieves higher ratings for expressiveness and conversational naturalness compared to existing studio baselines. A VITS2 model trained on Dialogs demonstrates the corpus's potential for expressive, dialog-like text-to-speech synthesis.
Why it matters: Dialogs provides a high-quality resource for developing more natural and expressive Russian dialog assistants, addressing a gap in available conversational speech data.