audio.cpp 0.5 Released: Dramatic TTS and Cross-lingual Voice Transfer
Acceptable-Cycle4645 · reddit · 2026-08-01
The local audio inference framework audio.cpp has released version 0.5, significantly expanding its model ecosystem and platform compatibility.
Core Model Updates:
- DramaBox: A new expressive TTS model based on the LTX-2.3 architecture. It allows prompt-directed control over emotions, pauses, laughs, and speaker behavior, achieving prompt-directed voice acting.
- Confucius4-TTS: Supports cross-lingual voice transfer, extracting a reference voice and synthesizing it in another language.
- Other Integrations: Added RVC for voice conversion, BS-RoFormer for vocal separation, and ASR models like Fun-ASR-Nano and Parakeet-TDT.
Platform & Engineering Optimizations:
- Introduced early HIP/ROCm support for AMD GPUs and accelerated Metal performance on Apple Silicon.
- Improved server and streaming paths, supporting live PCM ingest and cleaner streaming transcript deltas. The author called for community contributions in model performance optimization and lightweight WebUI development.
More from coding & agent
- SGLang Supports Inkling-Small on Dual DGX Spark, Hits 24 tok/s — ying11231 · 2026-08-01
- AI Agents Integrate with Sentry Alerts for Automated Incident Triage — zeeg · 2026-08-01
- Task Marketplace for Agentic Workflows Built on x402 Protocol Surfaces — MurrLincoln · 2026-08-01
- Enterprise MCP Deployment Pain Points: Who Enforces Agent Tool Access at Runtime? — Common_Dream9420 · 2026-08-01
- Continual Harness: A Self-Improving Architecture for AI Agents — chijinML · 2026-08-01
- Read-Only by Design: An MCP Server for Secure Agent Secret Management — nabsha · 2026-08-01