Skip to content
worth noting Tools and apps

AWS has published a tutorial on deploying real-time text-to-speech (vLLM-Omni, Qwen3-TTS) on Amazon SageMaker AI

only one source so far

AWS has released a tutorial on deploying the Qwen3-TTS model via the vLLM-Omni Deep Learning Container on Amazon SageMaker AI for streaming, low-latency voice responses, including code and a Gradio demo application.

AWS has published the first part of a tutorial series describing the deployment of specialized runtime containers (Deep Learning Containers, DLC) for bidirectional speech streaming. This part shows how to deploy the Qwen3-TTS model using the AWS vLLM-Omni DLC on Amazon SageMaker AI so that the application can play the generated speech before the entire response is finished. Text is sent and audio is received over a single persistent bidirectional connection, which can be tried via the included Gradio application.

The vLLM-Omni project extends the well-known vLLM inference framework from text generation to also handle processing and generation of audio, images, and video. AWS DLC packages tracked versions of vLLM-Omni into custom images and adds the routing middleware needed for integration with SageMaker. DLC version v1.5 adds support for bidirectional streaming and the v1/audio/speech/stream route used in this tutorial. The connection functions as a fully duplex WebSocket carried over HTTP/2 on port 8443, where a SageMaker sidecar forwards the connection to the native WebSocket route inside the container.

The deployment uses so-called instance pools – a priority-ordered list of compatible instance types (ml.g6.xlarge, ml.g6e.xlarge, ml.g5.xlarge, ml.g4dn.xlarge), of which SageMaker always runs only one. However, quota needs to be available for all the listed types at once, because SageMaker verifies it when creating the endpoint. The tutorial runs in the US East (N. Virginia) region, and the Gradio component plays the generated audio in 24kHz PCM chunks.

This part follows an earlier post describing the input side of the voice pipeline (transcription of microphone audio by the Voxtral-Mini-4B Realtime model) and complements it with the output side featuring streaming TTS. According to the source, the second part of the series is to cover image and video generation using the same vLLM-Omni DLC. The rest of the article was not available.

What changed

Why it matters

The tutorial gives developers a ready-made, reproducible procedure for building a low-latency voice feature (speech starts playing before the entire response is generated) without having to work out the streaming protocol and integration with the inference runtime themselves. For companies planning voice assistants or accessibility tools, this shortens the time to deployment, but it must also be kept in mind that the actual operating cost depends on which instance from the specified pool SageMaker assigns at any given moment.

Two audiences, two different impacts

What this means

01

For individuals

A developer who needs to build a low-latency voice application gets a ready-made procedure and code for deploying streaming TTS on SageMaker AI, including configuration of a bidirectional WebSocket connection and instance selection.

What to do Clone the sample repository (03-features/bidirectional-streaming-vLLM-Omni) and try deploying Qwen3-TTS via the Gradio application.
More practical updates →
02

For a business

Companies building voice assistants, customer support, or accessibility tools can use this procedure to deploy streaming TTS on managed AWS infrastructure without developing their own streaming backend.

Development
What to decide Before deployment, verify the available quota for all instance types in the pool (ml.g6.xlarge, ml.g6e.xlarge, ml.g5.xlarge, ml.g4dn.xlarge), because SageMaker may select a fallback instance with a different hourly price.
More business impacts →
Amazon SageMaker Qwen3-TTS real-time voice streaming text-to-speech vLLM-Omni

Check the original

Event sources

only one source so far · 1 publisher, 0 independent. We count feeds from the same owner only once.

1
AWS Machine Learning Blog primary source · first detected Build real-time voice applications with vLLM-Omni on SageMaker AI – Part 1