audio.cpp cuts Higgs Audio VRAM by 48%, speeds HTDemucs and PocketTTS 2x+
Acceptable-Cycle4645 · reddit · 2026-10-08
The audio.cpp project shipped runtime-level optimizations for local audio inference, with no loss in parity or correctness:
- Higgs Audio TTS: peak VRAM down 48% to 6GB
- HTDemucs: 2.21× faster on CUDA (1.95× on Vulkan)
- PocketTTS: 2.23× faster on CPU
- Smaller gains across MOSS-TTS, Qwen3-TTS, IndexTTS2 and others (6–21% VRAM savings)
The project now supports 110+ audio model families and 190+ variants across CUDA, Vulkan, Metal, AMD/HIP and CPU. The WebUI adds an experimental generation history feature, and the team is recruiting frontend contributors.
Related event: audio.cpp update cuts Higgs Audio memory 48%, supports 110+ models(2 posts)→
More from Infra
- NVIDIA's NeMo-DCR cuts 1T-model weight sync from 87.5 min to 150s, 12-40x faster checkpoint transfer — dair_ai · 2026-10-09
- Zyphra Speeds MoE Expert Routing Communication 2.63x on AMD MI300X GPUs — QuentinAnthon15 · 2026-10-09
- SkyPilot AI Infra Meetup at SF TechWeek: GPU Scheduling and SGLang Inference — rseroter · 2026-10-09
- d1-omni-600M sorts your voice notes in ~160ms, fully in-browser on WebGPU — iamrobotbear · 2026-10-09
- Cloud company books Delta Forge Two for 15 years at ~$5B before a single customer — YvesMulkers · 2026-10-09
- PyTorch Conference to Feature Meta's TorchTPU Cross-Hardware Portability Keynote — PyTorch · 2026-10-09