audio.cpp 0.5 Released: Introduces Expressive TTS and Cross-lingual Voice Transfer

Acceptable-Cycle4645 · reddit · 2026-08-01

Local audio inference project audio.cpp has released version 0.5, highlighting the DramaBox expressive TTS model (allowing prompt-directed control over emotion, laughs, and pauses) and Confucius4-TTS for cross-lingual voice transfer.

The update also integrates 7 new models including RVC for voice conversion, BS-RoFormer for vocal separation, and GLM-TTS. On the platform side, it adds early HIP/ROCm support for AMD GPUs, faster Metal performance on Apple Silicon, and improved server paths with live PCM ingest and cleaner streaming transcript deltas. The author calls for community contributions in scoped performance optimization and a lightweight WebUI alternative.

Related event: audio.cpp 0.5 Released: Introduces Dramatic TTS and Cross-lingual Voice Conversion(2 posts)→

Original post →

More from Infra

Infra channel →