AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Evaluating Voice Cloning And Multilingual TTS With The Open TTS Leaderboard on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get tech for your team delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Hugging Face has introduced an Open TTS Leaderboard that compares text-to-speech models using automated measures of speech accuracy, generation speed and speaker similarity. Its results are intended to make comparisons faster, not to determine which voices sound most natural or which models listeners prefer.

Hugging Face has launched the Open TTS Leaderboard, a public system for comparing text-to-speech models by speech accuracy, generation speed and speaker similarity. The project says automated evaluations can be run in hours, giving developers a faster way to screen models, but stresses that the scores do not measure naturalness or listener preference.

The leaderboard compares generated speech with the written prompts using word error rate and character error rate. Those measures rely on transcripts produced by Qwen3 automatic speech recognition. Its speed tests report offline generation performance using inverse real-time factor on an H200 GPU, alongside streaming time-to-first-audio measurements on an H200 GPU and a CPU.

For voice cloning, the system adds a speaker-similarity score based on WavLM embeddings from generated audio and reference recordings. The default ranking uses macro-average word error rate across English splits from Seed TTS Eval and CV3 Eval. Users can select other languages and enable a voice-cloning view. For Chinese, Japanese and Korean, the leaderboard uses character error rate; for languages beyond English and Chinese, the supplied announcement says Seed TTS Eval lacks audio and results come from CV3 Eval alone.

Hugging Face names k2-fsa/OmniVoice, fishaudio/s2-pro and FunAudioLLM/Fun-CosyVoice3-0.5B-2512 among strong multilingual models. For English error rates, it lists hexgrad/Kokoro-82M, Supertone/supertonic-3 and fishaudio/s2-pro among leading systems. These are results under the leaderboard’s chosen tests, not a general judgment about which models sound best. The separate “Listen” tab lets users compare generated audio and submit preferences.

At a glance
announcementWhen: Announced in September 2026; the source…
The developmentHugging Face launched an Open TTS Leaderboard that measures accuracy, speed and voice similarity across text-to-speech models.
At a glance
announcementWhen: Announced in material dated September 3…
The developmentHugging Face has launched an open leaderboard that evaluates text-to-speech models across multiple languages using objective performance metrics and offers audio comparisons for community feedback.

Faster Screening for Voice Models

The release offers developers and researchers a repeatable first comparison in a field where model releases can outpace evaluation. Hugging Face says its Hub held more than 8,000 TTS models as of September 30, 2026, while existing comparisons were fragmented and slow to update. The project says its objective tests can take a couple of hours, compared with weeks for arena voting, though that is the project’s description of evaluation time rather than an independently reported benchmark.

The separate measurements may help teams narrow choices based on their needs. A voice-agent developer may care about time-to-first-audio, while a multilingual product team may prioritize error rates across languages, and a voice-cloning workflow may focus on speaker similarity. The scores can also reveal tradeoffs: accuracy, speed and resemblance to a reference voice are distinct qualities, and a strong result in one does not establish strength in the others.

Automated rankings may also broaden attention to open-weight models if they are easier to evaluate without hosting each system for a voting arena. Hugging Face says that, as of September 30, it counted 16 open-weight models among 92 on Artificial Analysis and saw a similar skew on Voice Arena. That count is attributed to Hugging Face; it does not by itself establish why the models are represented unevenly or how the new leaderboard will change their visibility.

Amazon

AI voice cloning device

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Automated Scores and Listener Arenas

Text-to-speech systems turn written text into spoken audio, but ranking them involves both measurable performance and subjective listening. Existing services cited by Hugging Face, including TTS Arena v2, Artificial Analysis and Voice Arena, use listener comparisons: people hear two outputs, vote for a preference, and votes contribute to rankings, often using Elo scores calculated with a Bradley–Terry model.

Those arenas capture human preferences directly, but collecting votes takes time, and operating an arena can require hosting the models being compared. Hugging Face argues that open-weight systems are underrepresented on some existing rankings, partly because of the work involved in hosting them and because commercial providers may have stronger incentives to seek placement. Those explanations are the project’s assessment, not a finding established by the leaderboard announcement.

The new system combines standardized datasets and automatic measurements with a listening feature rather than treating the approaches as substitutes. Automated scores support faster, repeatable comparisons; listening gives users a way to judge qualities that error rates and speaker embeddings do not capture.

“The Open TTS Leaderboard does not replace human preference ranking.”

— Hugging Face

Amazon

multilingual text-to-speech software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Limits of the Published Rankings

The supplied announcement does not include the full model list, score tables, evaluation sample sizes or uncertainty ranges for individual results. It also does not establish how closely the automated metrics track listener judgments across languages, accents or speaking styles. Error rates depend on an automatic speech recognizer, while speaker similarity estimates preservation of voice identity; neither directly measures naturalness, expressiveness or preference.

Language coverage also varies by dataset: the announcement says some non-English scores rely on CV3 Eval alone. It is not yet clear how frequently Hugging Face will refresh rankings, how model or dataset changes will be handled, or what level of feedback would be needed before community votes affect the results. The named leaders should therefore be read as results within specific tests, not as a complete verdict on real-world quality.

Amazon

voice synthesis speaker

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Listening Votes and Future Updates

Users can visit the “Listen” tab, select a language and dataset, choose whether to compare voice cloning, and vote on audio outputs. Hugging Face asks users to sign in with an account to help limit spam and bot submissions. The project says it may incorporate community votes as feedback accumulates, but has not specified a threshold or timetable.

Further updates could clarify the score tables, evaluation coverage and refresh schedule, as well as whether listener preferences will be added to rankings. Until then, the leaderboard provides a faster way to compare selected technical measures, while users evaluating a model for a particular use case will still need to listen to its outputs and test it under relevant conditions.

Amazon

speech recognition and TTS tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What does the Open TTS Leaderboard measure?

It reports word and character error rates, offline generation speed, streaming time-to-first-audio and, in its voice-cloning view, speaker similarity. These measures cover different aspects of system performance.

Does the leaderboard show which model sounds most natural?

No. Hugging Face says the automated scores do not replace human judgments of naturalness or listener preference. The “Listen” tab lets users compare audio and submit preferences, but the announcement does not say when votes might affect rankings.

Which models does Hugging Face identify as strong?

For multilingual performance, the announcement names OmniVoice, fishaudio/s2-pro and Fun-CosyVoice3. For English error rates, it lists Kokoro-82M, supertonic-3 and fishaudio/s2-pro. These examples refer to the leaderboard’s evaluated measures, not an overall sound-quality ranking.

How quickly can an evaluation run?

Hugging Face says its objective-metric evaluation can take a couple of hours, compared with weeks for arena voting. The announcement does not provide an independent test of that timing estimate.

How will community votes affect future rankings?

The project says votes may be incorporated as feedback accumulates, and asks participants to sign in to reduce spam and bot activity. It has not announced a voting threshold or implementation schedule.

Primary source: Hugging Face · via ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Which Of The 9 Best AI Meeting Transcription Tools For 2026 Fits Your Workflow?

A 2026 roundup compares nine AI meeting recorders by workflow, transcription, app quality, battery life and subscription costs.

ElevenLabs, TwelveLabs, ThirteenLabs

Three AI companies—ElevenLabs, TwelveLabs, and ThirteenLabs—announced new advancements in voice and video synthesis, signaling rapid industry growth.

AI-Powered Chatbot Platforms: A Halloween Guide

Learn how AI-powered chatbot platforms work, what to test before buying, and how to protect users while measuring real results.

The Future Of Audio AI: How Grok Voice Realtime And xAI Lead The Way

xAI’s Grok Voice Realtime emerges as a key audio-to-audio model, signaling advances in real-time voice interaction technology, though details remain undisclosed.