TontaubeV1 Released: 2.9B Parameter Open TTS for Local Long-form Generation

EAVDR · reddit · 2026-08-28

The Tontaube team released TontaubeV1, a 2.9B-parameter open-weight TTS model focused on expressive speech, long-form generation, and low-latency local inference. It supports English and German, with zero-shot voice cloning. Technically, it uses four autoregressive models for codec generation and supports streaming unbounded long-form generation via a rolling context window. On an RTX 5090, it achieves 0.08 RTF for single text and as low as 0.02 RTF with batching. The current release requires at least 24GB VRAM, with quantized versions planned for smaller devices. Internal benchmarks show it scores 50.1% against ElevenLabs Flash v2.5 on prosody.

Original post →

More from Multimodal

Multimodal channel →