Struggles of Running Local TTS Inference on AMD GPUs in Windows

aboutthednm · reddit · 2026-08-05

A Windows user with an AMD 7900 XTX GPU shares the ecosystem struggles of deploying local Text-to-Speech (TTS) models. Despite testing several open-source projects like KokoroTTS and CosyVoice, achieving proper GPU acceleration (Vulkan/ROCm) on the AMD platform is difficult, and the prosody of current models feels flat.

The author outlines their ideal local TTS stack: GPU acceleration, a good selection of voices, an OpenAI-compatible endpoint, faster-than-real-time generation, and fitting within 23GB of VRAM. They are seeking advice from the community on better Windows + AMD setups.

Original post →

More from Infra

Infra channel →