🔍 Read the full analysis: Evaluating Voice Cloning And Multilingual TTS With The Open TTS Leaderboard on ThorstenMeyerAI.com
Get tech for your team delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Hugging Face has introduced an Open TTS Leaderboard that compares text-to-speech models using automated measures of speech accuracy, generation speed and speaker similarity. Its results are intended to make comparisons faster, not to determine which voices sound most natural or which models listeners prefer.
Hugging Face has launched the Open TTS Leaderboard, a public system for comparing text-to-speech models by speech accuracy, generation speed and speaker similarity. The project says automated evaluations can be run in hours, giving developers a faster way to screen models, but stresses that the scores do not measure naturalness or listener preference.
The leaderboard compares generated speech with the written prompts using word error rate and character error rate. Those measures rely on transcripts produced by Qwen3 automatic speech recognition. Its speed tests report offline generation performance using inverse real-time factor on an H200 GPU, alongside streaming time-to-first-audio measurements on an H200 GPU and a CPU.
For voice cloning, the system adds a speaker-similarity score based on WavLM embeddings from generated audio and reference recordings. The default ranking uses macro-average word error rate across English splits from Seed TTS Eval and CV3 Eval. Users can select other languages and enable a voice-cloning view. For Chinese, Japanese and Korean, the leaderboard uses character error rate; for languages beyond English and Chinese, the supplied announcement says Seed TTS Eval lacks audio and results come from CV3 Eval alone.
Hugging Face names k2-fsa/OmniVoice, fishaudio/s2-pro and FunAudioLLM/Fun-CosyVoice3-0.5B-2512 among strong multilingual models. For English error rates, it lists hexgrad/Kokoro-82M, Supertone/supertonic-3 and fishaudio/s2-pro among leading systems. These are results under the leaderboard’s chosen tests, not a general judgment about which models sound best. The separate “Listen” tab lets users compare generated audio and submit preferences.
Faster Screening for Voice Models
The release offers developers and researchers a repeatable first comparison in a field where model releases can outpace evaluation. Hugging Face says its Hub held more than 8,000 TTS models as of September 30, 2026, while existing comparisons were fragmented and slow to update. The project says its objective tests can take a couple of hours, compared with weeks for arena voting, though that is the project’s description of evaluation time rather than an independently reported benchmark.
The separate measurements may help teams narrow choices based on their needs. A voice-agent developer may care about time-to-first-audio, while a multilingual product team may prioritize error rates across languages, and a voice-cloning workflow may focus on speaker similarity. The scores can also reveal tradeoffs: accuracy, speed and resemblance to a reference voice are distinct qualities, and a strong result in one does not establish strength in the others.
Automated rankings may also broaden attention to open-weight models if they are easier to evaluate without hosting each system for a voting arena. Hugging Face says that, as of September 30, it counted 16 open-weight models among 92 on Artificial Analysis and saw a similar skew on Voice Arena. That count is attributed to Hugging Face; it does not by itself establish why the models are represented unevenly or how the new leaderboard will change their visibility.
As an affiliate, we earn on qualifying purchases.
Automated Scores and Listener Arenas
Text-to-speech systems turn written text into spoken audio, but ranking them involves both measurable performance and subjective listening. Existing services cited by Hugging Face, including TTS Arena v2, Artificial Analysis and Voice Arena, use listener comparisons: people hear two outputs, vote for a preference, and votes contribute to rankings, often using Elo scores calculated with a Bradley–Terry model.
Those arenas capture human preferences directly, but collecting votes takes time, and operating an arena can require hosting the models being compared. Hugging Face argues that open-weight systems are underrepresented on some existing rankings, partly because of the work involved in hosting them and because commercial providers may have stronger incentives to seek placement. Those explanations are the project’s assessment, not a finding established by the leaderboard announcement.
The new system combines standardized datasets and automatic measurements with a listening feature rather than treating the approaches as substitutes. Automated scores support faster, repeatable comparisons; listening gives users a way to judge qualities that error rates and speaker embeddings do not capture.
“The Open TTS Leaderboard does not replace human preference ranking.”
— Hugging Face
multilingual text-to-speech software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Limits of the Published Rankings
The supplied announcement does not include the full model list, score tables, evaluation sample sizes or uncertainty ranges for individual results. It also does not establish how closely the automated metrics track listener judgments across languages, accents or speaking styles. Error rates depend on an automatic speech recognizer, while speaker similarity estimates preservation of voice identity; neither directly measures naturalness, expressiveness or preference.
Language coverage also varies by dataset: the announcement says some non-English scores rely on CV3 Eval alone. It is not yet clear how frequently Hugging Face will refresh rankings, how model or dataset changes will be handled, or what level of feedback would be needed before community votes affect the results. The named leaders should therefore be read as results within specific tests, not as a complete verdict on real-world quality.
As an affiliate, we earn on qualifying purchases.
Listening Votes and Future Updates
Users can visit the “Listen” tab, select a language and dataset, choose whether to compare voice cloning, and vote on audio outputs. Hugging Face asks users to sign in with an account to help limit spam and bot submissions. The project says it may incorporate community votes as feedback accumulates, but has not specified a threshold or timetable.
Further updates could clarify the score tables, evaluation coverage and refresh schedule, as well as whether listener preferences will be added to rankings. Until then, the leaderboard provides a faster way to compare selected technical measures, while users evaluating a model for a particular use case will still need to listen to its outputs and test it under relevant conditions.
As an affiliate, we earn on qualifying purchases.
Key Questions
What does the Open TTS Leaderboard measure?
It reports word and character error rates, offline generation speed, streaming time-to-first-audio and, in its voice-cloning view, speaker similarity. These measures cover different aspects of system performance.
Does the leaderboard show which model sounds most natural?
No. Hugging Face says the automated scores do not replace human judgments of naturalness or listener preference. The “Listen” tab lets users compare audio and submit preferences, but the announcement does not say when votes might affect rankings.
Which models does Hugging Face identify as strong?
For multilingual performance, the announcement names OmniVoice, fishaudio/s2-pro and Fun-CosyVoice3. For English error rates, it lists Kokoro-82M, supertonic-3 and fishaudio/s2-pro. These examples refer to the leaderboard’s evaluated measures, not an overall sound-quality ranking.
How quickly can an evaluation run?
Hugging Face says its objective-metric evaluation can take a couple of hours, compared with weeks for arena voting. The announcement does not provide an independent test of that timing estimate.
How will community votes affect future rankings?
The project says votes may be incorporated as feedback accumulates, and asks participants to sign in to reduce spam and bot activity. It has not announced a voting threshold or implementation schedule.
Primary source: Hugging Face · via ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
