audio.cpp implements MiniMax-H3 for TTS, voice clone, music gen; 3x realtime on RTX 5090

Acceptable-Cycle4645 · reddit · 2026-08-17

The audio.cpp project implements MiniMax-H3's text-to-audio pipeline, enabling TTS, voice cloning, and music generation. It's more flexible than dedicated audio models, achieving up to 3x realtime on RTX 5090. It supports multi-speaker demos and simplifies DiT setup. It can also generate video frames. MiniMax-Music3 preview is available with CUDA/Vulkan/HIP support.

Related event: audio.cpp 0.6 Released with MiniMax-H3 Support(2 posts)→

Original post →

More from Multimodal

Multimodal channel →