audio.cpp 0.5 Released: Introduces Expressive TTS and Cross-lingual Voice Transfer
Acceptable-Cycle4645 · reddit · 2026-08-01
Local audio inference project audio.cpp has released version 0.5, highlighting the DramaBox expressive TTS model (allowing prompt-directed control over emotion, laughs, and pauses) and Confucius4-TTS for cross-lingual voice transfer.
The update also integrates 7 new models including RVC for voice conversion, BS-RoFormer for vocal separation, and GLM-TTS. On the platform side, it adds early HIP/ROCm support for AMD GPUs, faster Metal performance on Apple Silicon, and improved server paths with live PCM ingest and cleaner streaming transcript deltas. The author calls for community contributions in scoped performance optimization and a lightweight WebUI alternative.
More from Infra
- SGLang Supports Inkling-Small on Dual DGX Spark, Hits 24 tok/s — ying11231 · 2026-08-01
- macmon: Open-Source Terminal Performance Monitor for Apple Silicon Hits 1.8k Stars — tom_doerr · 2026-08-01
- Running Ideogram 4 Locally on Apple Silicon: Workflows, Memory Costs & JSON Prompts — DaLyon92x · 2026-08-01
- Scaling Kimi K3 on H200s: Engineering Insights from 1000+ Chips — hsu_byron · 2026-08-01
- DeepSeek Hits 40 tok/s Locally on M3 Ultra Mac Studio — zephyr_z9 · 2026-08-01
- Laguna Doubles Performance: Significant Mac Inference Speedup Without Speculative Decoding — gajesh · 2026-08-01