
What happened
Alibaba's Qwen released Qwen-Audio-3.1, five models for speech recognition, text-to-speech and real-time interaction, and cut prices: TTS about 70 percent, Realtime roughly 85 percent, ASR up to 95 percent.
Why it matters
The five-model lineup and the price cuts together mean speech recognition, text-to-speech and real-time interaction are now cheaper and more capable from one vendor, which could reshape what developers choose to build with.
What to watch
The price cuts are steep, but the real test is whether the five models deliver on multilingual and dialect recognition across real-world audio. Watch for user reports on Qwen Cloud and the linked blog.
WHO IT HITSDevelopers and product teams building voice features — such as call-center transcription or interactive voice assistants — may find Qwen-Audio-3.1's five models and lower prices a more affordable option for speech recognition, text-to-speech and real-time interaction.
Summaries like this, in your inbox every morning.
Alibaba's Qwen team has expanded its audio model lineup with Qwen-Audio-3.1, a set of five models spanning speech recognition (ASR), text-to-speech (TTS), and real-time interaction. The release includes improvements such as multilingual and dialect recognition in the ASR model, automatic cleanup of filler words, and new capabilities in ASR-Next for multi-speaker identification with timestamps and detection of emotions, ambient sounds, and machine noise. TTS now handles multilingual synthesis with natural cross-language voice transfer, and users can control emotion, speed, and style through simple text prompts.
The lineup also introduces TTS-Next, which pairs a language model with a diffusion approach to generate voice, sound effects, and background audio in a single pass, and a real-time model that supports simultaneous speaking and listening with instant interruption. According to Qwen, when the real-time model detects a low mood, it responds more slowly and with more empathy.
Alongside the release, Alibaba is cutting prices. TTS drops about 70 percent, Realtime roughly 85 percent, and ASR up to 95 percent. The combination of a broader, more capable lineup and steeper discounts could lower the cost for developers building speech-based applications. The outcome hinges on whether the models perform reliably across diverse languages and audio conditions in real-world use, and how quickly developers adopt the cheaper options. More details are available on the blog and on Qwen Cloud.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, LINE, or Slack.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Rabbit Inc. unveiled OS3, a cloud personal AI agent installed via a single command on up to five machines, wor…
At Snapdragon Summit 2026, Qualcomm CEO Cristiano Amon hosted Google SVP Rick Osterloh to discuss the past yea…

On September 21, SoftBank launched more than $11 billion in dollar and euro bonds ahead of a $10 billion OpenA…

Advanced Micro Devices crossed a $1 trillion market value for the first time on September 21 after shares jump…

Redis launched LangCache, a managed semantic cache that stores full question-response pairs outside the model…

ChatGPT Voice now runs on OpenAI's new GPT-6 Astra, Sol, and Luna models and can access plugins like email, ca…
