TTS silently dropped 17% of a passage, so this dev built VoxStage, a local multi-voice audiobook tool
HoujunDev · reddit · 2026-10-07
VoxStage is a new open-source (AGPL-3.0), fully local script-to-voice workstation for Apple Silicon Macs: paste unlabeled prose, get a multi-voice reading you can audition line by line, fix, redo, and export — no account, cloud API, or telemetry.
Origin: In an earlier voice-cloning test, a fluent-sounding passage silently lost 40 characters (17%) from the middle. This set two rules: generate sentence by sentence, and transcribe every line back with whisper.cpp, diffing against the script and flagging mismatches for human review — never auto-correcting.
Stack: Qwen3-TTS on MLX (preset/designed/cloned voices), llama.cpp local LLM (Qwen3-14B or Qwen3-30B-A3B on 32GB) drafting speaker attribution, whisper.cpp read-back checks.
Measured (M2 Max 32GB, Pride and Prejudice ch. 1): speaker draft for 35 units in 14.8s; 28 lines → 144.7s of audio in 51s (RTF 0.358); read-back check 28s.
Limits: speaker drafts usually need at least one human fix (review is the product); Chinese and English only; dev-style install (30 min, mostly downloads).
Also: single-sentence regeneration, audio-timed SRT/VTT subtitles, FCP7 XML for DaVinci Resolve, book/chapter cast management.
More from Multimodal
- Mirage's Tesseract + Opus 5.5 generated its entire launch video, exported to After Effects — aziz4ai · 2026-10-07
- Nano Banana 2.1 vs Midjourney 8.2: same-prompt image comparison surfaces — miilesus · 2026-10-07
- A lot of AI slop films are just porn, observer points out — moonsandhues · 2026-10-07
- First AI-Made Film Gets Traditional Theatrical Run, Submits for Best Animated Feature Oscar — Uncanny_Harry · 2026-10-07
- AI turns kids' doodles into fluffy living creatures — horned tiger included — anthara_ai · 2026-10-07
- Seedance 2.5 AI video of Dragon Ball battle wows with detail and motion — anthara_ai · 2026-10-07