Cartesia voice clone review: 10-second sample stays stable on long audio
dr_cintas · x · 2026-09-27
The author tested AI voice tools for long-form audio and cloned his own voice to read a full tutorial script. Most tools sounded good for the first few seconds but drifted in tone over time; Cartesia held up. It needs only 10 seconds of voice to create a clone, accepts up to 60 seconds to preserve accent, and he jokes the clone sounds "more real than I do."
Related event: Cartesia Wins Voice Cloning Test, Stays Stable on Long Audio(2 posts)→
More from Multimodal
- Claude Opus builds motion design entirely in code: Canvas, 20-sample motion blur, Node synth — mishig25 · 2026-09-27
- Qwen-Image 2.1 runs natively on M1 Max: ~24s at 512x512, FP16 no quantization — adel_b · 2026-09-27
- The AI-generated Nike commercial we made — Klitolovac · 2026-09-27
- RedFlag AI flags manipulation signals by analyzing voice, face and speech across 13 metrics — Perfect_Contest6714 · 2026-09-27
- URN Queue Media Viewer: open-source ComfyUI extension visualizes queued jobs — nobody123d2 · 2026-09-27
- Qwen-based consistent character workflow achieves ~99% character fidelity — Acceptable-Work8202 · 2026-09-27