
TontaubeV1 is a new open-weight TTS model for expressive, long-form speech.
It uses character-level tokenization and DualCodec to improve quality.
The model is trained on 200k hours and supports voice cloning.
What happened
Two developers released TontaubeV1, a 2.9B-parameter open-weight TTS model for expressive speech and long-form narration. It supports zero-shot voice cloning from up to one minute of reference audio and targets English and German.
Why it matters
The model uses character-level tokenization, which the authors found works better than the original BPE tokenizer in early experiments. This choice, plus building on the DualCodec multi-codebook audio codec, aims to improve long-form generation and low-latency local inference.
What to watch
The model was trained on 7 languages and about 200k hours of audio, but is mostly tested in English and German. Availability of the open weights and the specific licensing or access details were not mentioned in the announcement.
Ask the AI about this article →
The release highlights a deliberate design choice: character-level tokenization over the more common BPE approach in modern TTS models. The developers note that this generally led to better performance in their early tests, and they built the model on DualCodec, a multi-codebook audio codec, to handle the audio representation. These choices seem aimed at addressing common pain points in TTS—expressive delivery and long-form consistency—while keeping inference low-latency for local use.
Given the open-weight status, the model invites community testing and comparison with other TTS systems. The focus on English and German, despite training on 7 languages, suggests a priority on quality in those two markets. The lack of availability details in the announcement leaves room for users to seek out the official release channels for access and usage terms.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
AT&T, Dell Technologies, and AMD have announced OTel 2.0, the largest and best-performing open-source model bu…

Google released Gemini 3.7 Flash, its latest developer model, and unveiled the new Pixel 11 series

Top AI-native open source projects like Flue and tldraw are refusing external pull requests, often because the…

Anthropic has released its latest large language models, Claude Fable 5.1 and Claude Mythos 5.1

Phonely Ltd. launched Alma, a large language AI model built for voice agents and trained on over 10 million re…
Hugging Face released @huggingface/kernels, a library for running optimized WebGPU kernels from the Hugging Fa…
