
What happened
Hugging Face launched the Open TTS Leaderboard, ranking models by word error rate (WER/CER), H200 GPU speed measures, and voice-cloning speaker similarity, cutting evaluation from a couple weeks to a couple hours.
Why it matters
This gives open-source TTS models a standard, scalable way to be compared, which may help address their underrepresentation on human-vote arena leaderboards.
What to watch
The leaderboard relies on objective metrics like WER and speaker similarity, which may not fully capture naturalness or listener preference. Watch for community votes and feedback to shape future evaluations.
WHO IT HITSThis affects researchers and developers evaluating open-source TTS models, as well as voice-app builders who need to select models for multilingual or voice-cloning use cases. It may also help arena operators decide which open models to include.
Summaries like this, in your inbox every morning.
Hugging Face's Open TTS Leaderboard arrives as the pace of open-source text-to-speech releases has been incredible, with more than 8K TTS models on the Hugging Face Hub as of Sep 30, 2026. Evaluation, however, hasn't kept pace: it remains fragmented and unstandardized. Human preference scores like MOS or MUSHRA are the gold standard, and arena-based leaderboards such as TTS Arena v2, Artificial Analysis, and Voice Arena have served as reference points by collecting votes and computing Elo scores. But arenas cannot scale to keep up with releases, which may partly explain why open-source models are underrepresented there—only 16 of the 92 models on Artificial Analysis are open-weights, with a similar skew on Voice Arena.
The new leaderboard aims to fill that gap by using objective metrics to evaluate models on complementary aspects: intelligibility via word/character error rate, speed via RTFx and TTFA, and speaker similarity via cosine similarity of WavLM embeddings. This drops evaluation time from a couple weeks to a couple hours. It doesn't replace human preference ranking, as ASR-based WER is a proxy for intelligibility and speaker similarity estimates voice identity preservation, but neither directly measures naturalness or expressiveness. Still, these metrics can inform voting-based leaderboards which models to include.
The leaderboard is intended to be shaped by the community, with evaluation scripts to be open-sourced soon. Its success may hinge on whether the community contributes feedback and votes, and whether the objective metrics prove useful for selecting models in multilingual and voice-cloning scenarios. For now, it focuses on open-source models and multilingual performance, since English performance is not a suitable proxy for other languages.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Deepseek released open-source programming tools for Huawei's Ascend chips, including libraries for computation…

The Tokyo District Court dismissed Kenjiro Tsuda's demand that TikTok remove AI voice-clone videos, citing pri…

Anthropic published a September 29, 2026 review finding Z.ai's downloadable GLM-5.3 built working exploits in…

A writer published a Frog and Toad–style explainer about HuggingFace and the METR report, aimed at readers lik…

DeepL opened general availability of real-time speech-to-speech translation in DeepL Voice, supported by a new…

A Zenn handbook titled 生命科学研究 × AI 活用ハンドブック splits AI use into three levels — chat-only (near-zero learning co…
