Qwen3-TTS 1.7B hits 1.6x real-time voice cloning on CPU via llama.cpp

alexcovo_eth · x · 2026-09-11

A developer benchmarked Qwen3-TTS-12Hz-1.7B VoiceDesign on mainline llama.cpp (Q4KM): on an i7-12700H CPU it generated 3.44s of studio-grade cloned audio in 2.13s (1.61x real-time, zero VRAM, 8GB RAM); on an RTX 4060 Laptop it hit 3.22x real-time with only 3.5GB VRAM peak. No reference audio or fine-tuning needed — you describe age, gender, accent, mic proximity and emotion in plain text, and the zero-shot cloning penalty is nearly nonexistent.

Original post →

More from Infra

Infra channel →