Reddit user says LTX 2.3 audio beats TTS at breaths, laughs and intonation shifts
cptrios · reddit · 2026-07-23
A Reddit user says classic TTS and voice-cloning workflows still fail on realism details like breaths, sighs, complex intonation shifts, laughs, and groans.
They report that LTX 2.3’s audio engine is much better at generating those sounds, and that an ID-LoRa workflow can produce convincing laughs and other vocal effects in a reference voice.
The user is asking whether anyone has built an LTX / ID-LoRa speech-to-speech workflow that can take a custom voice recording and convert it into another reference voice, while preserving the realistic audio behaviors they care about.
More from Multimodal
- FLUX.2 Klein Drifts Hard on Character Expressions While Free Gemini Holds Likeness — wacomlover · 2026-09-11
- Tencent Hunyuan releases AuK code and weights on GitHub with ComfyUI and fine-tuning support — aigclink · 2026-09-11
- Creator turns Bahamut vs Tiamat rivalry into an AI cinematic battle with Midjourney, GPT Image 2 and Seedance — azed_ai · 2026-09-11
- invideo launches AI agent-powered editor to automate repetitive editing tasks — azed_ai · 2026-09-11
- fable 5.1 recreates The Starry Night with 256,157 JavaScript brush strokes — cedric_chee · 2026-09-11
- GPT-6 Astra + Hyper3D Rodin MCP Generates 3D Assets in One Agent Flow — ahuja_priyank · 2026-09-11