audio.cpp 0.4 adds 5 speech models and reports up to 37% VRAM savings
Acceptable-Cycle4645 · reddit · 2026-07-24
audio.cpp 0.4 adds several new speech models and makes GGUF a first-class format across the project.
- New support includes Higgs Audio v3 TTS 4B, Fish Audio S2 Pro, Voxtral Realtime ASR, plus community ports OuteTTS and VieNeu-TTS-v3.
- The project now supports 35 model families, and every released family now has GGUF packages.
- Benchmarks on an RTX 5090 suggest Q8 GGUF can deliver real gains: Higgs Audio TTS runs about 8.8×–10.1× real time on warmed requests, Fish Audio S2 Pro about 3.1×–3.4×, and Voxtral ASR about 15.7× with streaming TTFT around 171 ms.
- Compared with 16-bit GGUF, Q8 can be up to 1.5× faster and cut peak VRAM by up to 37%, though quality remains model-dependent.
Related event: audio.cpp 0.4 Adds Voice Models and Boosts VRAM Efficiency(2 posts)→
More from Infra
- NVIDIA says Nemotron 3 Ultra hit 97.1% on agentic RTL chip-design tasks — NVIDIAAI · 2026-07-27
- Local Qwen models power a robot that tests 78 smartphones’ battery life — gappyvalley · 2026-07-27
- MiniBot 2.40 adds xAI, HF Studio and vLLM support with inline media tools — Creative-Type9411 · 2026-07-27
- Apple smart glasses, Nvidia-SK AI data center deal, and Ctrip’s RMB 5.179 billion fine headline a tech roundup — APPSO · 2026-07-27
- DeepSeek funding rumor, EU AI transparency rules and OpenAI agent incident make a packed AI news roundup — 创业邦 · 2026-07-27
- QuixiCore argues native quantized kernels beat dequant-then-generic execution — QuixiAI · 2026-07-27