DIY full audiobook with local TTS: Higgs Audio beats VibeVoice, OmniVoice and Piper
HoodWinked69 · reddit · 2026-09-09
The author produced a complete audiobook of Schopenhauer's The Wisdom of Life using local TTS models, sharing workflows and the final audio.
- Tested VibeVoice 1.5B/7B, OmniVoice, Piper and Higgs Audio; Higgs is the favorite, OmniVoice the fastest
- Most models run via the ttsaudiosuite custom node pack in ComfyUI; workflows and a custom glow node pack are attached
- Hardware: RTX 5060 Ti 16GB + 32GB RAM, 2–4 min render per 1 min audio; detailed ComfyUI startup args (SageAttention 2.2.0, async-offload, etc.) with a warning not to blindly copy them
- Lessons: 15k–20k character chunks work well; future work includes automated text prep, reference cleanup, and automatic prosody insertion for consistent tone across chunks
More from coding & agent
- Driving tldraw Flash with Astra computer use to make a 2-minute animated movie — max__drake · 2026-09-09
- Yoav Goldberg proposes a new agent eval: locating real GitHub artifacts via API — yoavgo · 2026-09-09
- OpenAI's Navier-Stokes proof reportedly used ~10,000 coordinating AI agents — mark_k · 2026-09-09
- Open-source Aurora: Go LLM gateway fork claims 55x LiteLLM speed and fixes OpenCode session lockout — entitybtw · 2026-09-09
- Can agents resolve vague human references like 'the change that broke token refresh'? — yoavgo · 2026-09-09
- Dev open-sources local prompt optimizer on DSPy+Ollama, boosting llama3.2 3B from 39% to 89% with GEPA — MortisAndTen · 2026-09-09