audio.cpp 0.4 adds Higgs Audio v3, Fish Audio S2 Pro and GGUF support across 35 model families
Acceptable-Cycle4645 · reddit · 2026-07-24
audio.cpp 0.4 adds new TTS/ASR models and makes GGUF a first-class path
The project’s 0.4 release expands support with Higgs Audio v3 TTS 4B, Fish Audio S2 Pro, Voxtral Realtime ASR, plus community ports such as OuteTTS and VieNeu-TTS-v3.
Key changes and measurements:
- 35 model families are now supported.
- All released families now support GGUF.
- Ready-to-use GGUF packages are available, and Q8 is showing real speed and memory gains.
- On an RTX 5090, the author reports:
- Higgs Audio TTS: 8.8x–10.1x real-time after warmup; longform at 8.5x real-time
- Fish Audio S2 Pro: 3.1x–3.4x real-time warmed; 3.3x for longform
- Voxtral ASR: 15.7x real-time offline, with 171 ms streaming TTFT
- Compared with 16-bit GGUF, Q8 can be up to 1.5x faster and cut peak VRAM by up to 37%, depending on route and model.
The author notes Q8 is not universally better, so the release keeps the support matrix and performance report visible. There’s also a dedicated community-model area for ports that are useful and runnable even if still maturing.
Related event: audio.cpp 0.4 Adds Voice Models and Boosts VRAM Efficiency(2 posts)→
More from Infra
- OpenRouter agents now out-consume humans as AI usage arrives in three waves — AccBalanced · 2026-09-11
- Nvidia Is Now Core to Every Major Robotaxi Stack at Commercial Scale — pdamodaran · 2026-09-11
- 12 KV Cache Reduction Techniques Every AI Engineer Should Understand, Explained — blaizedsouza · 2026-09-11
- The shadow GPU capacity market is formalizing, with Meta selling excess compute to outside buyers — DavidLinthicum · 2026-09-11
- Engram's random reads don't suit SSDs; CPU-memory over NVLink could serve all 72 GPUs — bookwormengr · 2026-09-11
- 80% of the DIY LLM inference hype posters have already quit — it's brutally hard systems work — abhijithneil · 2026-09-11