Open-source TontaubeV1 TTS Model Uses Character-Level Tokenization
EAVDR · reddit · 2026-09-01
TontaubeV1 is a 2.9B-parameter open-weight TTS model built on Qwen3-1.7B, focusing on expressive speech, long-form narration, and low-latency local inference. It supports English and German with zero-shot voice cloning.
The release highlights two key design choices:
- Character-level Tokenization: Forces the model to process text as a sequence of individual characters instead of BPE tokens. This reduces out-of-distribution errors and rare token sequences while retaining language understanding capabilities.
- Chunking and Position Scheme: Processes text, semantic audio, and acoustic codebooks as parallel rows in a flat sequence. It assigns logical position IDs to align tokens representing the same audio frame across codebooks, using paired split markers and a sliding window to handle long-form context without leaking positions.
More from Models
- Testing unreleased Gemini 3.8 Flash: no citations shown for top-of-funnel queries — gaganghotra_ · 2026-09-03
- Hidden-bug eval across 105 issues: Fable 5.1 finds 43, none fixes all — cost per model compared — PawelHuryn · 2026-09-03
- X open-sources new For You algorithm code: long dwell drives retrieval, bots can trigger account review — Kyrannio · 2026-09-03
- Gemini 3.8 Flash reverse-engineers Kerbal save files to build and land a Mun rocket — dosco · 2026-09-03
- User calls out model for double-standard answers on gendered scenario questions — Ribbitz_bow_tie27 · 2026-09-03
- Redditor Predicts Astra Model Release Tomorrow at 1pm PT Based on X Teasers — dolo937 · 2026-09-03