TTS silently dropped 17% of a passage, so this dev built VoxStage, a local multi-voice audiobook tool

HoujunDev · reddit · 2026-10-07

VoxStage is a new open-source (AGPL-3.0), fully local script-to-voice workstation for Apple Silicon Macs: paste unlabeled prose, get a multi-voice reading you can audition line by line, fix, redo, and export — no account, cloud API, or telemetry.

Origin: In an earlier voice-cloning test, a fluent-sounding passage silently lost 40 characters (17%) from the middle. This set two rules: generate sentence by sentence, and transcribe every line back with whisper.cpp, diffing against the script and flagging mismatches for human review — never auto-correcting.

Stack: Qwen3-TTS on MLX (preset/designed/cloned voices), llama.cpp local LLM (Qwen3-14B or Qwen3-30B-A3B on 32GB) drafting speaker attribution, whisper.cpp read-back checks.

Measured (M2 Max 32GB, Pride and Prejudice ch. 1): speaker draft for 35 units in 14.8s; 28 lines → 144.7s of audio in 51s (RTF 0.358); read-back check 28s.

Limits: speaker drafts usually need at least one human fix (review is the product); Chinese and English only; dev-style install (30 min, mostly downloads).

Also: single-sentence regeneration, audio-timed SRT/VTT subtitles, FCP7 XML for DaVinci Resolve, book/chapter cast management.

Original post →

More from Multimodal

Multimodal channel →