faster-qwen3-tts brings quantized Qwen3-TTS to Apple Silicon via GGML
andimarafioti · x · 2026-09-25
The open-source faster-qwen3-tts project now supports Apple Silicon via GGML, enabling real-time local speech generation with quantized Qwen3-TTS weights on Macs. It offers streaming and non-streaming modes: the Torch backend uses CUDA Graph (NVIDIA GPU required), while the GGML backend supports CUDA and Metal (macOS 14+). Install with pip; native Apple Silicon Python auto-selects the Metal runtime. 1.4k stars on GitHub.
More from Infra
- Theoretical Neuroscience podcast: how neuromorphic computing could speed up AI and cut energy use — neurovium · 2026-09-25
- Goldman: Big Tech AI capex to jump 50%+ to $1.2T in 2027, $1.4T in 2028 — rohanpaul_ai · 2026-09-25
- Robot simulation is the cleanest anti-gaming incentive, argues Bittensor SN49 — bittingthembits · 2026-09-25
- SemiAnalysis: NVIDIA B200 serving DeepSeek can yield up to $15B annual profit per gigawatt — zephyr_z9 · 2026-09-25
- Qdrant open-sources Supernova embedding tool and 10B FineWeb dataset — qdrant_engine · 2026-09-25
- Founder argues datacenter buildouts are a dead end — AI models will shrink like mainframes did — draginol · 2026-09-25