Qwen3-TTS 1.7B hits 1.6x real-time voice cloning on CPU via llama.cpp
alexcovo_eth · x · 2026-09-11
A developer benchmarked Qwen3-TTS-12Hz-1.7B VoiceDesign on mainline llama.cpp (Q4KM): on an i7-12700H CPU it generated 3.44s of studio-grade cloned audio in 2.13s (1.61x real-time, zero VRAM, 8GB RAM); on an RTX 4060 Laptop it hit 3.22x real-time with only 3.5GB VRAM peak. No reference audio or fine-tuning needed — you describe age, gender, accent, mic proximity and emotion in plain text, and the zero-shot cloning penalty is nearly nonexistent.
More from Infra
- Keep the Claude Desktop Workflow, Swap in Local Models via Ollama for Privacy — Technovangelist · 2026-09-11
- LithosAI ships Day-0 API inference for DeepSeek-V4.1-Flash at 250+ tokens/s per user — JiaZhihao · 2026-09-11
- 1:26 continuous aerial AI video made entirely on a Mac with MiniMax H3 — cocktailpeanut · 2026-09-11
- KV cache gets QAT too: why this model beats others at fp4 KV cache — stochasticchasm · 2026-09-11
- Commentary: Anthropic loads shift to Google plus AWS slice, OpenAI doubles down on Azure — ericwdolan · 2026-09-11
- Nvidia claims Vera Rubin delivers 50X throughput per MW and 35X lower token cost vs Blackwell Ultra — Beth_Kindig · 2026-09-11