Open-Source Real-Time Streaming TTS Model Gepard 1.0 Released

ylankgz · reddit · 2026-07-08

The team has open-sourced Gepard 1.0 (Apache 2.0), a real-time streaming TTS model designed for real-time conversational scenarios. It features a streaming-first design—generating audio frame-by-frame as text arrives, without waiting for the full sentence. It has roughly 555M parameters (Qwen3.5 0.8B backbone with 14 layers + Nemo NanoCodec, 22.05kHz). On a single RTX 5090, it achieves about 20x real-time rate and ~50ms time-to-first-audio. A single RTX Pro 6000 Blackwell (96GB) supports up to 256 parallel streams. It supports zero-shot voice cloning with just a few seconds of reference audio and covers US/UK English, Mexican Spanish, Brazilian Portuguese, and Dutch. On the Seed-TTS-eval benchmark, compared to VoxCPM2, Fish-S2, OmniVoice, Qwen3-TTS, Echo-TTS, and Chatterbox Turbo, it leads in perceived quality NISQA-MOS (4.25) and is the cleanest in noise/coloration/discontinuation metrics. The honest trade-off is that the streaming-first design sacrifices speaker similarity (SIM 0.585) and WER (0.036).

Related event: Open-Source Real-Time Streaming TTS Model Gepard 1.0 Released(2 posts)→

Original post →

More from Multimodal

Multimodal channel →