AIToday
Large Language ModelsAudio & SpeechAI Business & IndustryTHE DECODERPublished: Sep 23, 2026, 22:00 JST

Alibaba's Qwen-Audio-3.1 cuts audio AI prices up to 95 percent

Alibaba's Qwen-Audio-3.1 cuts audio AI prices up to 95 percent

3 Key Points

  1. What happened

    Alibaba's Qwen released Qwen-Audio-3.1, five models for speech recognition, text-to-speech and real-time interaction, and cut prices: TTS about 70 percent, Realtime roughly 85 percent, ASR up to 95 percent.

  2. Why it matters

    The five-model lineup and the price cuts together mean speech recognition, text-to-speech and real-time interaction are now cheaper and more capable from one vendor, which could reshape what developers choose to build with.

  3. What to watch

    The price cuts are steep, but the real test is whether the five models deliver on multilingual and dialect recognition across real-world audio. Watch for user reports on Qwen Cloud and the linked blog.

WHO IT HITSDevelopers and product teams building voice features — such as call-center transcription or interactive voice assistants — may find Qwen-Audio-3.1's five models and lower prices a more affordable option for speech recognition, text-to-speech and real-time interaction.

Not sure about something? Ask the AI

Summaries like this, in your inbox every morning.

Context & Analysis

Alibaba's Qwen team has expanded its audio model lineup with Qwen-Audio-3.1, a set of five models spanning speech recognition (ASR), text-to-speech (TTS), and real-time interaction. The release includes improvements such as multilingual and dialect recognition in the ASR model, automatic cleanup of filler words, and new capabilities in ASR-Next for multi-speaker identification with timestamps and detection of emotions, ambient sounds, and machine noise. TTS now handles multilingual synthesis with natural cross-language voice transfer, and users can control emotion, speed, and style through simple text prompts.

The lineup also introduces TTS-Next, which pairs a language model with a diffusion approach to generate voice, sound effects, and background audio in a single pass, and a real-time model that supports simultaneous speaking and listening with instant interruption. According to Qwen, when the real-time model detects a low mood, it responds more slowly and with more empathy.

Alongside the release, Alibaba is cutting prices. TTS drops about 70 percent, Realtime roughly 85 percent, and ASR up to 95 percent. The combination of a broader, more capable lineup and steeper discounts could lower the cost for developers building speech-based applications. The outcome hinges on whether the models perform reliably across diverse languages and audio conditions in real-world use, and how quickly developers adopt the cheaper options. More details are available on the blog and on Qwen Cloud.

FAQ
What models are included in Qwen-Audio-3.1?
The lineup has five models covering speech recognition (ASR), text-to-speech (TTS), and real-time interaction.
How much did Alibaba cut prices for these audio models?
TTS drops about 70 percent, Realtime roughly 85 percent, and ASR up to 95 percent.
What new features does the ASR model add?
The ASR model improves multilingual and dialect recognition and automatically cleans up filler words and repetitions. ASR-Next adds multi-speaker identification with timestamps and detects emotions, ambient sounds, and machine noise.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Rabbit unveils OS3, a cloud AI agent for up to five machinesSiliconANGLE AI · 1h ago
  • Qualcomm, Google widen Snapdragon Summit 2026 tie-upDIGITIMES Asia · 1h ago
  • Redis LangCache: 15x faster, 70% cheaper on cache hitsDaily Dose of Data Science · 1h ago

AI-summarized, only the topics you pick — one digest a day via Email, LINE, or Slack.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleMcKinsey's Catlin: same AI agent task can cost 30x more per run