AWS Tutorial: Stream Qwen3-TTS Speech on SageMaker via vLLM-Omni Bidirectional Streaming
AWS ML Blog · rss · 2026-09-29
AWS ML Blog's Part 1 tutorial deploys Qwen3-TTS on SageMaker AI using the vLLM-Omni Deep Learning Container:
- Architecture: SageMaker bidirectional streaming over an HTTP/2 WebSocket (port 8443); an inference sidecar forwards to vLLM-Omni's native v1/audio/speech/stream route — text in, 24 kHz PCM audio chunks out over one persistent connection, so playback starts before generation finishes
- vLLM-Omni DLC: extends vLLM to text/audio/image/video with heterogeneous autoregressive+diffusion pipelines and OpenAI-compatible APIs; v1.5 adds the bidirectional streaming label and routing middleware
- Instance pools: priority list from ml.g6.xlarge with fallbacks to g6e/g5/g4dn; quota needed per entry and hourly cost varies by selected type
- Steps: clone the aws-samples repo, set the execution role (Python 3.12+, boto3≥1.40.0, HTTP/2 client≥0.4.0), run deploybidistream.py (smoke test saves validation-output.wav), then launch the Gradio client to stream speech with chosen voice/language
- The series continues with WhisperX and llama.cpp; Part 2 covers image/video generation
More from Infra
- Kernel Design Agents optimize Kimi Delta Attention kernels, up to 2.96x speedup — songhan_mit · 2026-09-29
- OpenRouter inference providers slash GLM 5.3 output price from $4.40 to $1.61 in a month — AccBalanced · 2026-09-29
- ASCI Q supercomputer crashed every 6.5 hours until a neutron beam exposed cosmic rays as the culprit — lauriewired · 2026-09-29
- Dual 5070+5060Ti local LLM setup too slow: is a single RX 7900XTX the fix? — rawdikrik · 2026-09-29
- How SNI actually works: the TLS feature powering CDNs and multi-tenant routing — HankYeomans · 2026-09-29
- Well-known compute broker joins Compute Exchange as senior deals lead — ns123abc · 2026-09-29