IndexTTS 2.5 Launches: 2.28x Faster Inference and Zero-Shot Cross-Lingual Emotion Transfer
Xianbao_QIAN · x · 2026-08-12
IndexTTS 2.5 has been released, featuring broader language coverage, zero-shot emotion transfer, and a 2.28x improvement in inference speed.
According to its technical report, the core improvements include:
- Semantic Codec Compression: Reduces the frame rate from 50 Hz to 25 Hz, halving the sequence length and significantly lowering training and inference costs.
- Architectural Upgrade: Replaces the U-DiT backbone in the S2M module with a more efficient Zipformer architecture, achieving parameter reduction and faster mel-spectrogram generation.
- Multilingual Extension: Introduces three cross-lingual strategies (boundary-aware alignment, token-level concatenation, instruction-guided generation) supporting Chinese, English, Japanese, Spanish, and Arabic. It enables robust emotion transfer without target-language emotional training data.
- Reinforcement Learning Optimization: Applies Group Relative Policy Optimization (GRPO) in the post-training of the T2S module to improve pronunciation accuracy and naturalness.
More from Multimodal
- Using Narrative Logic to Guide AI Image Generation: Advanced Midjourney Prompting — tisch_eins · 2026-08-12
- AI Music Analyzer Showcased: Auto-Extracts Song Highlights — threepointone · 2026-08-12
- MiniMax Video Drift: Close-Ups Lose Identity After 3 Seconds — ItsMilaVoss · 2026-08-12
- LTX 2.5 Tested: Generates 15s 1080P Video in 8 Minutes on RTX 4080 — skyrimer3d · 2026-08-12
- MiniMax H3 local full-body shots suffer face distortion, likely due to VRAM limits — Loud-Guitar1920 · 2026-08-12
- How Cultural Differences Shape AI Video Model Preferences: H3 vs. LTX 2.5 — xbobos · 2026-08-12