AIToday
Audio & SpeechOpen-Source AIHugging Face BlogPublished: Oct 1, 2026, 01:00 JST

Hugging Face Open TTS Leaderboard cuts eval from weeks to hours

Hugging Face Open TTS Leaderboard cuts eval from weeks to hours

3 Key Points

  1. What happened

    Hugging Face launched the Open TTS Leaderboard, ranking models by word error rate (WER/CER), H200 GPU speed measures, and voice-cloning speaker similarity, cutting evaluation from a couple weeks to a couple hours.

  2. Why it matters

    This gives open-source TTS models a standard, scalable way to be compared, which may help address their underrepresentation on human-vote arena leaderboards.

  3. What to watch

    The leaderboard relies on objective metrics like WER and speaker similarity, which may not fully capture naturalness or listener preference. Watch for community votes and feedback to shape future evaluations.

WHO IT HITSThis affects researchers and developers evaluating open-source TTS models, as well as voice-app builders who need to select models for multilingual or voice-cloning use cases. It may also help arena operators decide which open models to include.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

Hugging Face's Open TTS Leaderboard arrives as the pace of open-source text-to-speech releases has been incredible, with more than 8K TTS models on the Hugging Face Hub as of Sep 30, 2026. Evaluation, however, hasn't kept pace: it remains fragmented and unstandardized. Human preference scores like MOS or MUSHRA are the gold standard, and arena-based leaderboards such as TTS Arena v2, Artificial Analysis, and Voice Arena have served as reference points by collecting votes and computing Elo scores. But arenas cannot scale to keep up with releases, which may partly explain why open-source models are underrepresented there—only 16 of the 92 models on Artificial Analysis are open-weights, with a similar skew on Voice Arena.

The new leaderboard aims to fill that gap by using objective metrics to evaluate models on complementary aspects: intelligibility via word/character error rate, speed via RTFx and TTFA, and speaker similarity via cosine similarity of WavLM embeddings. This drops evaluation time from a couple weeks to a couple hours. It doesn't replace human preference ranking, as ASR-based WER is a proxy for intelligibility and speaker similarity estimates voice identity preservation, but neither directly measures naturalness or expressiveness. Still, these metrics can inform voting-based leaderboards which models to include.

The leaderboard is intended to be shaped by the community, with evaluation scripts to be open-sourced soon. Its success may hinge on whether the community contributes feedback and votes, and whether the objective metrics prove useful for selecting models in multilingual and voice-cloning scenarios. For now, it focuses on open-source models and multilingual performance, since English performance is not a suitable proxy for other languages.

FAQ
How does the Open TTS Leaderboard evaluate models?
It uses objective metrics: word/character error rate (WER and CER) for intelligibility, inverse real-time factor (RTFx) and time-to-first-audio (TTFA) for speed on an H200 GPU and CPU, and cosine similarity (SIM) of WavLM speaker embeddings for voice cloning.
Which models lead on English WER?
hexgrad/Kokoro-82M, Supertone/supertonic-3, and fishaudio/s2-pro lead on English WER when averaged on Seed TTS Eval and CV3 Eval splits.
Can I compare model outputs directly?
Yes, the Listen tab lets you compare generated outputs and vote, but you need to log in with an HF account to prevent spam.
Hugging Face BlogRead Original Article

AI news that matters for your work, in one minute a day

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleAmazon Bedrock Knowledge Bases hits 90.5% claim retrieval recall