AIToday
Large Language ModelsAudio & SpeechAmazon AI BlogPublished: Sep 29, 2026, 04:00 JST

AWS streams Qwen3-TTS audio via vLLM-Omni on SageMaker AI

AWS streams Qwen3-TTS audio via vLLM-Omni on SageMaker AI

3 Key Points

  1. What happened

    AWS published Part 1 of a tutorial that deploys Qwen3-TTS on Amazon SageMaker AI using the vLLM-Omni Deep Learning Container, streaming 24 kHz PCM audio chunks over one HTTP/2 WebSocket.

  2. Why it matters

    Because Qwen3-TTS returns audio before finishing the full response, voice agents can apparently start playing speech sooner, cutting the silent pause that makes interactive voice applications feel slow.

  3. What to watch

    The sample uses SageMaker instance pools that try ml.g6.xlarge first and fall back to ml.g6e.xlarge, ml.g5.xlarge, or ml.g4dn.xlarge, but you need quota for every listed type and the hourly cost can change with the instance selected.

WHO IT HITSAI application developers and platform engineers who build voice agents, accessibility tools, or customer service assistants on AWS can now follow a concrete sample to stream text-to-speech output. Teams running GPU endpoints on SageMaker AI will also need to manage endpoint quota and cleanup, as a running GPU endpoint continues to incur charges.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The tutorial builds on an earlier AWS post that covered the input side of a voice pipeline: streaming microphone audio to the Voxtral-Mini-4B Realtime speech-to-text model and returning transcription events. This new post adds the output side, sending text to Qwen3-TTS and streaming generated speech back through the vLLM-Omni DLC. Together, the two examples show complementary paths: stream microphone audio into transcription, then stream the application's text response back as speech. The orchestration between those two endpoints remains outside the walkthrough.

The vLLM-Omni project extends vLLM beyond text generation to serve models that process or generate text, audio, images, and video, and its heterogeneous pipeline abstraction coordinates multi-stage workflows including autoregressive and diffusion stages. AWS packages tracked vLLM-Omni releases in its own images and adds routing middleware for SageMaker AI. The post uses Qwen3-TTS as a focused example because streamed speech provides a direct way to demonstrate bidirectional audio output, and it notes that the same DLC family will be applied to image and video generation in Part 2.

For teams building voice applications, the practical question is whether this pattern lowers the silent pauses that make voice agents feel slow. The sample's reliance on instance pools and quota for each included type suggests the approach's usefulness may hinge on whether capacity and cost line up with the intended workload.

FAQ
What model does the tutorial deploy?
It deploys Qwen3-TTS, a text-to-speech model, through the AWS vLLM-Omni Deep Learning Container on Amazon SageMaker AI.
What audio format does the streaming output use?
The sample configures pulse-code modulation (PCM) output and plays each 24 kHz PCM chunk as the Gradio client receives it.
What happens if SageMaker cannot get the first instance type?
The sample uses instance pools: SageMaker tries ml.g6.xlarge first, then falls back to ml.g6e.xlarge, ml.g5.xlarge, or ml.g4dn.xlarge when capacity is unavailable.
Amazon AI BlogRead Original Article

AI news that matters for your work, in one minute a day

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleLenfest expands AI fellowships with $5 million from OpenAI