Cohere has launched Transcribe Arabic, a state-of-the-art, enterprise-ready speech recognition model for Arabic speakers. The model is available as open source and is designed to capture the full diversity of spoken Arabic.
Why it matters: This release addresses the need for accurate transcription across diverse Arabic dialects, with open-source availability enabling broader enterprise and developer adoption.
Alpha Bank has expanded its collaboration with ElevenLabs by using ElevenAgents to build a new service aimed at simplifying customer communication and providing faster, more accessible interactions. The announcement was made on ElevenLabs' blog.
Why it matters: This partnership highlights the increasing use of AI agents in banking to improve customer service.
ElevenLabs has introduced ElevenAgents Spotlight, an observation and improvement layer for its ElevenAgents platform. The tool is designed to help increase resolution rates and customer satisfaction across all channels.
Why it matters: This enhancement offers a systematic approach to improving AI agent performance, which could advance customer service automation.
Fyxer, a meeting notetaker, uses ElevenLabs' Scribe v2 speech-to-text model, resulting in a 15% relative lift in user conversion. The integration demonstrates the model's effectiveness for real-time transcription.
Why it matters: This case study highlights the tangible benefits of advanced speech-to-text models for user engagement in productivity tools.
Italian telecom CoopVoce has deployed ElevenLabs' voice AI for customer support calls, resulting in a 10% increase in customer willingness to engage with its AI assistant within days. The AI assistant uses natural-sounding speech to handle customer inquiries, aiming to improve the user experience.
Why it matters: This deployment demonstrates how advanced voice AI can quickly enhance customer engagement in telecom support.
ElevenLabs has launched new tools on its ElevenMusic platform, allowing users to record vocals, add musical ideas, or transform completed tracks. These features are designed to help users create, reshape, and evolve their music.
Why it matters: This development expands ElevenLabs' AI capabilities into music production, providing creators with new ways to generate and modify audio content.
Robotics engineer Binh Pham used the Allen Institute for AI's MolmoAct 2 to build a voice-controlled robot that won the South Park Commons embodied AI hackathon. This achievement highlights the capabilities of open models in advancing robotics innovation.
Why it matters: This demonstrates the potential of open models like MolmoAct 2 to accelerate progress in embodied AI and robotics.
Google DeepMind has integrated its most advanced music generation model, Lyria 3, into the Gemini app. Users can now create 30-second tracks using text or images.
Why it matters: This development makes AI-powered music creation more accessible to a broad audience through a widely used app.
Mistral AI has announced Voxtral, a speech transcription model that transcribes audio at the speed of sound. The model is intended for real-time transcription applications.
Why it matters: Voxtral could advance real-time speech recognition by enabling faster transcription services.
Warner Music Group and Stability AI have announced a collaboration to develop responsible AI tools for music creation. The partnership aims to combine WMG's advocacy for principled innovation with Stability AI's expertise in commercially-safe generative audio.
Why it matters: This partnership highlights a major label's commitment to integrating AI into music production with a focus on ethical and legal safeguards.
Universal Music Group and Stability AI have announced a strategic alliance to co-develop next-generation professional music creation tools. These tools will use responsibly trained generative AI and are intended to support the creative process for artists, producers, and songwriters.
Why it matters: This partnership highlights a major move toward integrating generative AI into professional music creation with an emphasis on responsible development.
Stability AI has launched Stable Audio 2.5, its latest audio model developed specifically for enterprise-grade sound production. The model introduces advancements in quality and control, enabling dynamic compositions that can be tailored to custom brand requirements.
Why it matters: This release represents the first audio model designed for enterprise use, supporting scalable and customizable sound production for brands.
Stability AI, in partnership with Arm, has open-sourced Stable Audio Open Small, a compact variant of its text-to-audio model. The new model is designed to be smaller and faster while maintaining output quality and prompt adherence, enabling real-world deployment on devices.
Why it matters: This collaboration enables high-quality AI audio generation on smartphones and other edge devices, expanding accessibility and potential on-device applications.
A new study revisits XAI-guided adaptive fusion (XGAF) for multimodal emotion and sentiment recognition, using TreeSHAP attribution magnitudes to weight unimodal and cross-modal experts. On MELD 7-class emotion recognition, sum-abs XGAF nearly matches early fusion (0.5983 vs 0.6018) and significantly outperforms late fusion (0.4598). On CMU-MOSEI 3-class sentiment, sum-abs XGAF slightly exceeds early fusion (0.6519 vs 0.6485).
Why it matters: This work provides a transparent empirical analysis of how SHAP reduction methods and expert dimensionality affect modular multimodal fusion, offering a principled alternative to monolithic early fusion.
OpenAI has launched GPT-Live, a new feature that allows ChatGPT to talk, listen, and formulate answers simultaneously. This update is designed to make conversations with ChatGPT more natural and fluid.
Why it matters: The improvement could make AI interactions feel more human-like, enhancing user experience with conversational agents.
Paris-based AI voice startup Gradium has raised a $100 million seed round backed by Nvidia. The company plans to use the funds to open a Bay Area office and compete for talent, aiming to strengthen its position in the global AI ecosystem.
Why it matters: This large seed round from a major AI hardware company signals strong investor confidence in AI voice technology and the globalization of AI startups.
OpenAI has announced GPT-Live, a new generation of voice models designed for natural human-AI interaction. This model now powers ChatGPT Voice, enhancing real-time conversational capabilities.
Why it matters: GPT-Live represents a significant advancement in voice AI, enabling more fluid and natural spoken interactions with AI systems.
Apple Machine Learning Research has published a study on Text-to-Sounding-Video (T2SV) generation, which aims to produce videos with synchronized audio from text. The research identifies challenges such as text conditioning bottlenecks and unclear cross-modal fusion mechanisms, and proposes solutions to improve alignment between modalities.
Why it matters: This work advances multimodal AI by addressing the synchronization of video and audio from text, which has applications in content creation and accessibility.
Hugging Face and Cerebras have partnered to enable real-time voice AI using the Gemma 4 model. The collaboration utilizes Cerebras hardware to achieve low-latency inference for voice applications.
Why it matters: This partnership could make real-time conversational AI more practical by reducing latency in voice AI systems.
Hugging Face has introduced the FFASR Leaderboard, a new benchmark designed to evaluate automatic speech recognition (ASR) systems under real-world conditions. The leaderboard aims to provide a more practical assessment of ASR performance beyond standard datasets.
Why it matters: This benchmark addresses the gap between lab-tested ASR accuracy and real-world performance, helping developers choose models that work reliably in diverse acoustic environments.