Hugging Face launches Open TTS Leaderboard with objective metrics for TTS models
Hugging Face has launched the Open TTS Leaderboard, evaluating text-to-speech models with objective metrics (WER, CER, RTFx, TTFA, speaker similarity) instead of community voting. According to the company, evaluation now takes hours instead of weeks and better accounts for open-source models.
Hugging Face has launched the Open TTS Leaderboard, a ranking for evaluating text-to-speech models using objective metrics instead of community voting. According to the company, the current evaluation of TTS models is fragmented and relies mainly on arena-style leaderboards (e.g. Voice Arena, Artificial Analysis), where users compare the outputs of two models and the resulting Elo score is calculated using the Bradley–Terry model. According to the company, this approach cannot keep up with the pace of new TTS model releases and leads to underrepresentation of open-source models — as of 30 September 2026, Artificial Analysis lists only 16 out of 92 models with open weights, and a similar ratio applies on Voice Arena. Another problem with arena leaderboards, according to the company, is the inconsistency of raters over time.
The new leaderboard measures intelligibility using word/character error rate (WER, CER) between the input text and the transcript of the generated audio (via the Qwen3 ASR model), speed using RTFx (batch inference) and TTFA (time to first audio chunk during streaming, on H200 GPU as well as CPU), and voice similarity using cosine similarity of WavLM embeddings between the generated audio and the reference recording. According to the company, this approach shortens model evaluation from weeks to hours. At the same time, the company states that objective metrics do not replace human assessment of the naturalness and expressiveness of speech.
The leaderboard covers English, Chinese, Japanese, Korean and other languages, includes a separate view for voice cloning, a "Listen" tab for listening comparisons and community voting (after logging in with an HF account), and a "Streaming" tab ranked by TTFA. Among the models with the best English WER are listed Kokoro-82M, supertonic-3, and s2-pro; in the multilingual comparison, OmniVoice, s2-pro, and Fun-CosyVoice3-0.5B stand out; for streaming, pocket-tts is mentioned. For details, see the source article.
Why it matters
For developers of voice applications, the leaderboard offers a faster and more measurable way to choose a TTS model based on intelligibility, speed, and voice fidelity, including streaming latency data important for interactive voice agents. According to the company itself, however, the metrics do not replace listening-based assessment of speech naturalness, so human evaluation remains important for the final decision.
Two audiences, two different impacts
What this means
For individuals
Developers and creators of voice applications gain a tool for quickly comparing TTS models by intelligibility, speed, and voice similarity, without having to wait for the results of community voting.
For a business
Companies developing voice agents or TTS products can shorten the time needed to select a model from weeks (collecting votes in arenas) to hours thanks to objective metrics, and have access to streaming latency data (TTFA) important for interactive applications.
DevelopmentCheck the original
Event sources
only one source so far · 1 publisher, 0 independent. We count feeds from the same owner only once.