audio.cpp 0.4 adds 5 speech models and reports up to 37% VRAM savings
Acceptable-Cycle4645 · reddit · 2026-07-24
audio.cpp 0.4 adds several new speech models and makes GGUF a first-class format across the project.
- New support includes Higgs Audio v3 TTS 4B, Fish Audio S2 Pro, Voxtral Realtime ASR, plus community ports OuteTTS and VieNeu-TTS-v3.
- The project now supports 35 model families, and every released family now has GGUF packages.
- Benchmarks on an RTX 5090 suggest Q8 GGUF can deliver real gains: Higgs Audio TTS runs about 8.8×–10.1× real time on warmed requests, Fish Audio S2 Pro about 3.1×–3.4×, and Voxtral ASR about 15.7× with streaming TTFT around 171 ms.
- Compared with 16-bit GGUF, Q8 can be up to 1.5× faster and cut peak VRAM by up to 37%, though quality remains model-dependent.
Related event: audio.cpp 0.4 Adds Voice Models and Boosts VRAM Efficiency(2 posts)→
More from Infra
- LLM Serving Metrics Thread: Why TPOT and Uptime Make or Break User Experience — abhijithneil · 2026-09-11
- PlanetScale launches sharded Postgres: 768 servers acting as one, 1PB scale — dhruv2038 · 2026-09-11
- Can a 7900 XTX 24GB run Qwen locally? Reddit seeks ROCm tok/s benchmarks — thenomadexplorerlife · 2026-09-11
- RTK Terminal Compression Cuts Tokens but Leaves Your AI Coding Bill Unchanged — Bartaseth · 2026-09-11
- SF Compute founder: buying compute is 'an absolutely awful experience' right now — IgorCarron · 2026-09-11
- SmolVM open-sources persistent computer infrastructure for agents that outlive chat sessions — aniketmaurya · 2026-09-11