faster-qwen3-tts brings quantized Qwen3-TTS to Apple Silicon via GGML

andimarafioti · x · 2026-09-25

The open-source faster-qwen3-tts project now supports Apple Silicon via GGML, enabling real-time local speech generation with quantized Qwen3-TTS weights on Macs. It offers streaming and non-streaming modes: the Torch backend uses CUDA Graph (NVIDIA GPU required), while the GGML backend supports CUDA and Metal (macOS 14+). Install with pip; native Apple Silicon Python auto-selects the Metal runtime. 1.4k stars on GitHub.

Original post →

More from Infra

Infra channel →