
What happened
AWS published Part 1 of a tutorial that deploys Qwen3-TTS on Amazon SageMaker AI using the vLLM-Omni Deep Learning Container, streaming 24 kHz PCM audio chunks over one HTTP/2 WebSocket.
Why it matters
Because Qwen3-TTS returns audio before finishing the full response, voice agents can apparently start playing speech sooner, cutting the silent pause that makes interactive voice applications feel slow.
What to watch
The sample uses SageMaker instance pools that try ml.g6.xlarge first and fall back to ml.g6e.xlarge, ml.g5.xlarge, or ml.g4dn.xlarge, but you need quota for every listed type and the hourly cost can change with the instance selected.
WHO IT HITSAI application developers and platform engineers who build voice agents, accessibility tools, or customer service assistants on AWS can now follow a concrete sample to stream text-to-speech output. Teams running GPU endpoints on SageMaker AI will also need to manage endpoint quota and cleanup, as a running GPU endpoint continues to incur charges.
Summaries like this, in your inbox every morning.
The tutorial builds on an earlier AWS post that covered the input side of a voice pipeline: streaming microphone audio to the Voxtral-Mini-4B Realtime speech-to-text model and returning transcription events. This new post adds the output side, sending text to Qwen3-TTS and streaming generated speech back through the vLLM-Omni DLC. Together, the two examples show complementary paths: stream microphone audio into transcription, then stream the application's text response back as speech. The orchestration between those two endpoints remains outside the walkthrough.
The vLLM-Omni project extends vLLM beyond text generation to serve models that process or generate text, audio, images, and video, and its heterogeneous pipeline abstraction coordinates multi-stage workflows including autoregressive and diffusion stages. AWS packages tracked vLLM-Omni releases in its own images and adds routing middleware for SageMaker AI. The post uses Qwen3-TTS as a focused example because streamed speech provides a direct way to demonstrate bidirectional audio output, and it notes that the same DLC family will be applied to image and video generation in Part 2.
For teams building voice applications, the practical question is whether this pattern lowers the silent pauses that make voice agents feel slow. The sample's reliance on instance pools and quota for each included type suggests the approach's usefulness may hinge on whether capacity and cost line up with the intended workload.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Rosenblatt raised its Amazon price target to $360 from $335, kept a Buy rating, and called concerns that AI sh…

In Oracle's 80-turn evaluation, Oracle AI Agent Memory held input near 1,300 tokens per request while flat his…

Deepseek released open-source programming tools for Huawei's Ascend chips, including libraries for computation…

LinkedIn CEO Dan Shapero told the Wall Street Journal that job seekers have sent out 30% more applications tha…

Confluent's 2026 Data Streaming Report found just 17% of Japanese firms run agentic AI in production, the lowe…

Restate raised a $20 million Series A led by Singular, with Redpoint Ventures and Capital One Ventures, after…
