audio.cpp cuts Higgs Audio TTS VRAM by 48%, now supports 110+ audio model families
Acceptable-Cycle4645 · reddit · 2026-10-08
Local audio inference project audio.cpp shipped a batch of runtime-level optimizations with no quality compromise:
- Higgs Audio TTS: peak VRAM down 48% to under 6GB, the headline improvement
- HTDemucs: 2.21× faster on CUDA, 1.95× on Vulkan
- PocketTTS: 2.23× faster on CPU with 9% less RAM
- Others: MOSS-TTS cloning saves 21% VRAM, Qwen3-TTS 16-20%, IndexTTS2/2.5 12%, ACE-Step 1.06-1.20× faster
The WebUI adds an experimental generation history feature to revisit past outputs and restore their settings. The project now supports 110+ audio model families and 190+ variants across CUDA, Vulkan, Metal, AMD/HIP, and CPU, and is recruiting frontend contributors.
Related event: audio.cpp update cuts Higgs Audio memory 48%, supports 110+ models(2 posts)→
More from Infra
- Hobbyist runs 512K context locally on CPU/RAM/SSD/GPU hybrid at 1,526 tok/s prefill — HankYeomans · 2026-10-08
- ExecuTorch BoF session set for PyTorch Conference North America 2026 — PyTorch · 2026-10-08
- CMP 170HX unlock mod to lift 40GB cards to 48GB is progressing — chemist_slime · 2026-10-08
- AI engineer shares a skill checklist spanning KV cache, quantization, and agent guardrails — ghumare64 · 2026-10-08
- New video tutorial: use Hugging Face Buckets with Xet to upload only missing chunks — lhoestq · 2026-10-08
- Open training run: EDA beats GDN-2, auto expert failover survives 7 host crashes — jon_durbin · 2026-10-08