audio.cpp 0.4 adds Higgs Audio v3, Fish Audio S2 Pro and GGUF support across 35 model families
Acceptable-Cycle4645 · reddit · 2026-07-24
audio.cpp 0.4 adds new TTS/ASR models and makes GGUF a first-class path
The project’s 0.4 release expands support with Higgs Audio v3 TTS 4B, Fish Audio S2 Pro, Voxtral Realtime ASR, plus community ports such as OuteTTS and VieNeu-TTS-v3.
Key changes and measurements:
- 35 model families are now supported.
- All released families now support GGUF.
- Ready-to-use GGUF packages are available, and Q8 is showing real speed and memory gains.
- On an RTX 5090, the author reports:
- Higgs Audio TTS: 8.8x–10.1x real-time after warmup; longform at 8.5x real-time
- Fish Audio S2 Pro: 3.1x–3.4x real-time warmed; 3.3x for longform
- Voxtral ASR: 15.7x real-time offline, with 171 ms streaming TTFT
- Compared with 16-bit GGUF, Q8 can be up to 1.5x faster and cut peak VRAM by up to 37%, depending on route and model.
The author notes Q8 is not universally better, so the release keeps the support matrix and performance report visible. There’s also a dedicated community-model area for ports that are useful and runnable even if still maturing.
Related event: audio.cpp 0.4 Adds Voice Models and Boosts VRAM Efficiency(2 posts)→
More from Infra
- YC talk on BCI x AI says infrastructure is what really determines speed — garrytan · 2026-07-27
- A 13B model ran on a no-GPU PC by paging weights from SSD via llama.cpp — ID_R_McGregor · 2026-07-27
- llama.cpp warns that GGUFs made before a recent change must be regenerated — EconomySerious · 2026-07-27
- RTX 5090 local tests show Qwen Q6 can drop to 15 tok/s at 80k context — LFAdvice7984 · 2026-07-27
- Surprising Ubuntu Setup: NVIDIA 5090 PC Becomes the Easiest AI Rig — _xjdr · 2026-07-27
- TSMC reportedly plans 5%–10% price hikes in 2027 to cover rising costs — Beth_Kindig · 2026-07-27