Open TTS Leaderboard Grades 8K Models Without Arena Votes
By Indie Kings | October 10, 2026
Updated October 10, 2026: The Open TTS Leaderboard from Hugging Face ranks open text to speech models with objective intelligibility, speed, and speaker similarity tests instead of crowd votes alone, per the Hugging Face blog article Open TTS Leaderboard by Eric Bezzam, Steven Zheng, Eustache Le Bihan, and mrfakename dated Sep 30 2026. This guide explains the design, the metrics, and the default English ranking in plain terms.
Image: Open TTS leaderboard thumbnail. Credit: Hugging Face blog.
Why TTS evaluation was fragmented
There were more than 8K TTS models on the Hub as of Sep 30 2026, per the Hugging Face blog article Open TTS Leaderboard by Eric Bezzam, Steven Zheng, Eustache Le Bihan, and mrfakename. That count explains why choosing a model felt difficult for builders, creators, and researchers. A large catalog with no shared scoreboard leaves each team to run its own tests, trust vendor claims, or rely on scattered demos. The article frames that situation as fragmented evaluation, where results do not transfer cleanly from one report to the next.
Speech arenas offered one answer before this leaderboard, per the same Hugging Face blog article. The article names TTS Arena v2, Artificial Analysis, and Voice Arena as examples of arena style evaluation. Those systems collect pairwise votes from listeners, then convert the votes into rankings with methods such as Elo or Bradley Terry scoring. Pairwise voting can capture taste and preference in a direct way, since a listener hears two clips and picks a winner. The article treats that strength as real while also listing practical limits that motivated a different design.
One limit is coverage of open models, per the Hugging Face blog article. The article states that only 16 of 92 models on Artificial Analysis were open weights at the time of writing. That ratio matters for readers who want to download, host, inspect, or fine tune a model. A scoreboard dominated by closed systems gives less guidance to a builder who needs an open artifact with clear licensing and local control. The Open TTS Leaderboard answers that gap by focusing on open models that can be evaluated in a repeatable pipeline.
Another limit is hosting and incentives, per the Hugging Face blog article. The article explains that open models need hosting before they can appear in a vote based arena, while commercial providers seek placement because visibility has business value. That combination can skew which models appear and how often they are tested. A community model without a hosted demo may never collect enough votes, even if the weights are strong. A commercial model with placement support may collect many votes, even if the comparison pool is uneven.
Voter consistency is a further concern, per the Hugging Face blog article. Pairwise votes depend on who listens, what prompts are used, what audio setup is used, and what criteria each voter applies. One voter may reward clarity, another may reward warmth, and a third may reward speed of delivery. Consistency across many voters is hard to maintain over time. The article presents objective metrics as a complement that reduces this variance, because the same audio input and the same scoring code produce the same number on repeated runs.
This background sets up the leaderboard as a design choice rather than a claim that one model is universally superior. The default ranking reflects specific English intelligibility tests, per the Hugging Face blog article. Other toggles change the picture for multilingual use, voice cloning, and streaming. Readers should treat each view as an answer to a narrow question, not as a final verdict on quality. That framing runs through the full article and it shapes how the tables below should be read.
For Indie Kings readers, the practical takeaway is that model count alone does not equal progress unless measurement keeps pace. More than 8K TTS models on the Hub means discovery is the bottleneck, per the Hugging Face blog article. A shared pipeline that drops evaluation time from weeks to hours changes what teams can test in a normal work cycle. The next section explains the three metric families that make that faster loop possible.
What the leaderboard measures
The Open TTS Leaderboard uses objective metrics for intelligibility, speed, and speaker similarity, per the Hugging Face blog article. Intelligibility is measured with WER and CER through Qwen3 ASR, per the same source. Speed is measured with inverse RTF for batched offline synthesis on H200 plus time to first audio for streaming with batch size 1 on H200 and CPU, per the same source. Speaker similarity is measured with cosine SIM of WavLM embeddings, per the same source. Evaluation time drops from weeks to hours under this pipeline, per the same source.
Intelligibility deserves a careful explanation because it drives the default ranking. WER stands for word error rate and CER stands for character error rate in standard speech testing practice. The pipeline generates speech from text, transcribes it with Qwen3 ASR, then compares the transcript with the intended text. A lower error rate means the spoken output preserved more of the intended words or characters. That check is narrow by design, since it tests whether speech can be understood, not whether it sounds beautiful.
CER matters most for languages where word boundaries work differently than in English, per the Hugging Face blog article. The article notes that CJK uses CER, which covers Chinese, Japanese, and Korean text where character level scoring gives a more stable signal. A word based score can mislead when tokenization is ambiguous, while a character based score tracks insertions, deletions, and substitutions at a finer grain. Readers comparing multilingual toggles should remember that the unit of scoring changes with the script. That detail keeps cross language comparisons honest.
Speed is split into two views because offline and streaming use cases stress different parts of a system, per the Hugging Face blog article. Inverse RTF for batched offline synthesis on H200 reflects throughput when many utterances can be processed together without interaction. Time to first audio for streaming with batch size 1 on H200 and CPU reflects responsiveness when playback should start quickly. A model can score well in one view and less well in the other, depending on architecture, runtime, and batching behavior. Builders should pick the speed column that matches their product.
Offline batching on H200 gives a controlled hardware baseline for throughput, per the Hugging Face blog article. The article specifies H200 hardware for the batched offline path, which means results are comparable because the accelerator is held constant. Inverse RTF expresses how much audio is produced per unit of processing time, so a higher value points to faster bulk generation. That view suits audiobook rendering, dataset creation, batch narration, and other jobs where total completion time matters more than first byte latency. It does not answer how snappy a live voice agent will feel.
Streaming time to first audio answers the interactive question, per the Hugging Face blog article. The article specifies batch size 1 on H200 and CPU for this view, which mirrors a single user waiting for playback to begin. The Streaming tab ranks TTFA on 50 CV3 prompts with warm up dropped and median reported, per the same source. Dropping warm up removes cold start noise, while the median reduces the influence of rare outliers. The result is a responsiveness ranking that is easier to interpret than a single anecdotal test.
Speaker similarity uses cosine SIM of WavLM embeddings, per the Hugging Face blog article. In plain terms, the pipeline extracts speaker representations from generated audio and reference audio with WavLM, then measures cosine similarity between those vectors. A higher SIM value means the generated voice stayed closer to the reference identity in embedding space. That signal matters for voice cloning views, where preserving identity is part of the task. It does not measure naturalness on its own, which the article states directly.
The table below restates the metric families in one place, per the Hugging Face blog article. Standing is Per HF, meaning each row reflects the design as described by Hugging Face in that Sep 30 2026 article. Use this table when matching a product need to the right column, then read the ranking tables with that need in mind.
| Metric family | Score used | Hardware and setup | What it answers |
|---|---|---|---|
| Intelligibility | WER and CER via Qwen3 ASR | Transcription scoring pipeline | Did the output preserve the intended words or characters |
| Speed offline | Inverse RTF batched offline | H200 batched synthesis | How fast is bulk generation when prompts can be batched |
| Speed streaming | Time to first audio with batch size 1 | H200 and CPU streaming | How quickly playback can start for one interactive request |
| Speaker similarity | Cosine SIM of WavLM embeddings | Embedding comparison | How close the generated voice stayed to the reference identity |
| Evaluation cost | Pipeline runtime from weeks to hours | Shared open scripts | How quickly a team can repeat the full evaluation |
The metric design reflects a tradeoff that the article makes explicit. Objective scores are repeatable, fast, and comparable across many models, per the Hugging Face blog article. Human preference still captures qualities that numbers miss, especially naturalness and taste. The leaderboard therefore works best as a filter and a diagnostic tool, not as a replacement for listening. Readers should use WER and CER to shortlist intelligible models, use speed columns to match deployment constraints, and use SIM to check identity preservation, then listen before shipping.
Another practical point is that the eval scripts are open sourced, per the Hugging Face blog article. The article gives the repository name as github.com/huggingface/open-tts-leaderboard in plain text without a link in this guide, per the task rule for naming that path. Open scripts let outside teams inspect prompts, scoring logic, and hardware settings. That transparency is what allows the claim that evaluation drops from weeks to hours to be checked rather than taken on trust. Teams can rerun the same steps, adapt them to new prompts, or report issues against a shared baseline.
English default ranking and leaders
The default ranking is a macro average WER on English Seed TTS Eval plus CV3 Eval in zero shot mode, per the Hugging Face blog article. Macro averaging means each evaluation set contributes equally to the final order rather than letting a larger set dominate. Zero shot means the model generates speech without task specific tuning on those prompts, which tests generalization rather than memorization. English is the default view, while other languages and tasks use separate toggles. That scope matters because a leader in this view is not automatically a leader elsewhere.
The English leaders named in the Hugging Face blog article are hexgrad/Kokoro-82M, Supertone/supertonic-3, and fishaudio/s2-pro. The article presents those three as the top of the default English intelligibility ranking at the time of writing. This guide repeats that order only as the leaderboard design, not as an endorsement of any model rank beyond what is printed in the source. Model versions change, new submissions arrive, and reruns can shift positions. Readers should check the live board for the current order before making a selection.
Seed TTS Eval and CV3 Eval serve different roles in the default average, per the Hugging Face blog article. Seed TTS Eval focuses on English intelligibility under controlled prompts, while CV3 Eval adds a second English zero shot signal. Averaging the two reduces the chance that a model wins by fitting quirks of one set. A model that leads the macro average therefore showed strength across both English signals in the same pipeline. That breadth is more informative than a single set win.
The table below lists the default English leaders exactly as named in the Hugging Face blog article. Standing is Per HF, meaning the order and names reflect that Sep 30 2026 source and not independent testing by Indie Kings. Use the table as a starting point for shortlisting, then verify licensing, voice options, runtime needs, and output quality for your own prompts.
| Default English order | Model identifier | What the rank means Per HF |
|---|---|---|
| Leader group 1 | hexgrad/Kokoro-82M | Top English macro average WER group Per HF |
| Leader group 2 | Supertone/supertonic-3 | Top English macro average WER group Per HF |
| Leader group 3 | fishaudio/s2-pro | Top English macro average WER group Per HF |
| Evaluation sets | Seed TTS Eval plus CV3 Eval | Macro average WER in zero shot mode Per HF |
| Scope note | English default view only | Multilingual, cloning, and streaming use separate views Per HF |
It helps to state what a WER lead does and does not prove. A lead means the pipeline transcribed more words correctly from that model output under the stated English sets, per the Hugging Face blog article. It does not prove warmer prosody, better acting, more suitable pacing, or stronger multilingual skill. A narration project may care most about sustained clarity over long chapters, while a game project may care more about short line delivery and character fit. Both projects can start from the same WER shortlist, then diverge after listening tests.
Version sensitivity is another reason to treat the table as a snapshot. The identifiers above include specific model names and version style suffixes as printed in the Hugging Face blog article. A later checkpoint, a different sampling setting, or a changed voice preset could move results. The open scripts help here because they let teams rerun the same English sets on the exact checkpoint they plan to ship. That rerun habit matters more than memorizing a podium order.
Builders should also separate intelligibility from deployment fit. An English leader with strong WER still needs to meet latency, memory, license, and language coverage needs for the target product. Offline batch throughput on H200 does not guarantee smooth streaming on CPU, and English strength does not guarantee CJK strength. The leaderboard design anticipates this by placing toggles and tabs alongside the default view. The next sections explain those alternate views.
Multilingual toggles and voice cloning
Multilingual toggles extend the leaderboard beyond the English default, per the Hugging Face blog article. The article notes that CJK uses CER rather than WER, which keeps scoring stable for Chinese, Japanese, and Korean scripts. Separate toggles let readers inspect language specific behavior without mixing incompatible units into one average. A model that looks average in English may be strong in another language, and the reverse can also be true. That separation prevents misleading blended scores.
The strong multilingual models named in the Hugging Face blog article are k2-fsa/OmniVoice, fishaudio/s2-pro, and FunAudioLLM/Fun-CosyVoice3-0.5B-2512. The article presents those three as strong multilingual performers in the relevant toggles. This guide repeats those names only as the leaderboard design, not as an endorsement beyond what is printed. As with the English table, readers should verify the live board, the language list, and the checkpoint names before choosing. Language coverage can vary widely even within one model family.
Seeing fishaudio/s2-pro in both the English leader group and the strong multilingual group is informative but not decisive, per the Hugging Face blog article. It means that one identifier appeared among the leaders in two different views in that source. That breadth can simplify shortlists for teams that need English plus additional languages from one artifact. It still does not answer voice style, license terms, streaming latency, or cloning behavior. Each of those needs its own check in the matching tab or toggle.
The voice cloning toggle adds SIM to the evaluation, per the Hugging Face blog article. In that view, speaker similarity joins intelligibility as part of the comparison. The article states that some models improve with reference audio, naming bosonai/higgs-tts-3-4b and openbmb/VoxCPM2 as examples. Improvement with reference audio means the cloning input helped those models preserve identity or clarity relative to their non cloning behavior in the same framework. Readers testing custom voices should focus on this toggle rather than the default English ranking alone.
Reference audio quality remains important even when a model can use it well. A clean reference with one speaker, stable loudness, and minimal background noise gives the similarity check a clearer target. A noisy or multi speaker reference makes identity harder to preserve, regardless of architecture. The SIM score measures closeness in WavLM embedding space, not human approval, so listening remains necessary after the numbers narrow the field. Use SIM to detect identity drift, then use ears to judge fit.
Multilingual and cloning views also change how teams should plan test prompts. English narration prompts do not expose tone handling, script mixing, or name pronunciation in other languages. Cloning prompts without a consistent reference voice do not expose identity stability across sessions. Teams should therefore build small prompt packs that mirror real use, including short dialogue, long narration, names, numbers, and code switched phrases where relevant. The leaderboard toggles guide that process, but they do not replace product specific listening.
A final caution is that language toggles and cloning toggles answer different questions, so they should not be blended into one claim. A strong multilingual score does not imply strong cloning, and strong cloning with one reference voice does not imply every voice will clone equally well. The article keeps these views separate for that reason, per the Hugging Face blog article. Readers get cleaner decisions by matching one toggle to one requirement at a time.
Listen tab and Streaming tab
The Listen tab compares outputs and collects votes with HF login, per the Hugging Face blog article. That tab brings human judgment back into a pipeline that is otherwise fully objective. A visitor can play outputs from different models on the same prompts, compare them directly, and submit a preference vote after logging in. Login adds accountability relative to anonymous voting, though it does not remove taste differences. The article presents this tab as a complement to metric rankings rather than a replacement.
Direct comparison on shared prompts is the main value of the Listen tab. Metrics can say that two models have similar WER, yet listeners may strongly prefer one voice over the other. Prosody, pacing, warmth, crispness, and handling of punctuation often decide those preferences. Hearing the same sentence from several models makes those differences easier to notice than reading numbers alone. Teams choosing a customer facing voice should use this tab after metric filtering.
Vote collection also helps future analysis, per the Hugging Face blog article. Each vote adds a human preference signal tied to specific outputs and prompts. Over time, that data can reveal where objective scores and human taste agree and where they diverge. A model with strong WER but weak preference votes may sound robotic, flat, or mismatched to listener expectations. A model with weaker WER but strong votes may be expressive enough that listeners forgive minor errors. Both patterns are useful for research and product choice.
The Streaming tab ranks TTFA on 50 CV3 prompts with warm up dropped and median reported, per the Hugging Face blog article. TTFA stands for time to first audio, which is the delay before playback can begin. Fifty prompts give a broader sample than a single anecdote, dropping warm up removes cold start effects, and the median keeps one slow outlier from dominating the result. That method produces a responsiveness order that is easier to trust than ad hoc timing.
The article names kyutai/pocket-tts as streaming well on GPU and CPU, per the Hugging Face blog article. This guide repeats that observation only as the leaderboard design, not as an endorsement beyond what is printed. Streaming strength on both H200 and CPU matters because many interactive products cannot assume a flagship accelerator at serving time. A model that starts quickly on modest hardware can enable live agents, dialogue systems, and accessibility tools where bulk throughput alone would not help.
Builders should match the Streaming tab to interactive requirements and the offline speed column to batch requirements. A podcast renderer that processes chapters overnight cares more about inverse RTF in batched offline mode on H200, per the Hugging Face blog article. A voice agent that must answer within a conversational pause cares more about TTFA with batch size 1 on H200 and CPU, per the same source. Choosing the wrong speed metric can lead to a model that benchmarks well but feels slow in the real product.
Taken together, the two tabs show why the leaderboard keeps objective and subjective signals side by side. Metrics give speed, repeatability, and broad coverage across many models. Listening and voting capture preference, naturalness, and context fit. Streaming timing connects both worlds to deployment reality. Readers who use all three views will make fewer mistakes than readers who rely on one number alone.
What it does not replace
The leaderboard does not replace human preference, per the Hugging Face blog article. WER and SIM proxy intelligibility and identity, not naturalness, per the same source. That sentence is the most important scope limit in the whole design. A transcript match can confirm that words were preserved, and an embedding similarity can confirm that identity stayed close, but neither score hears warmth, emphasis, humor, or dramatic timing. Human listeners remain the final judge for those qualities.
Naturalness covers a wide range of traits that resist simple scoring. Sentence rhythm, breath placement, emphasis on key words, handling of questions and exclamations, and smooth transitions between sentences all shape whether speech feels pleasant over minutes or hours. Two models can share similar WER yet feel very different in a long audiobook chapter. The Listen tab exists for that reason, since direct playback reveals pacing and tone that error rates cannot encode.
Identity similarity has its own limits even when SIM is high. Cosine SIM of WavLM embeddings measures closeness in a learned representation space, per the Hugging Face blog article. That space is useful for tracking whether a cloned voice drifted toward a generic speaker or stayed near the reference. It does not guarantee that a specific listener will recognize the voice, nor that the voice suits a new role or language. Casting decisions still need human review with the actual script.
The article also keeps expectations realistic about arena voting. Pairwise votes plus Elo or Bradley Terry scoring can capture preference at scale, but the article lists hosting needs, placement incentives, and voter consistency as problems, per the Hugging Face blog article. The objective pipeline avoids those problems by using shared prompts, fixed hardware references, and open code. The tradeoff is that it cannot fully capture taste. Using both approaches together gives a fuller picture than either approach alone.
Open scripts support that balanced use, per the Hugging Face blog article. The repository name github.com/huggingface/open-tts-leaderboard appears here in plain text with no link, per the formatting rule for this guide. Teams can read the scoring logic, confirm prompt lists such as Seed TTS Eval and CV3 Eval, and rerun tests on their own shortlists. That audit path matters for production choices, because it turns a public ranking into a repeatable internal check. It also helps researchers propose better prompts or metrics without starting from zero.
For product teams, the recommended sequence is simple. Start with the default English macro average WER to filter for intelligibility, per the Hugging Face blog article. Switch to multilingual toggles when the audience needs more than English, remembering that CJK uses CER. Add the cloning toggle with SIM when custom voices matter, and note that bosonai/higgs-tts-3-4b and openbmb/VoxCPM2 improved with reference audio in the source. Check streaming TTFA on 50 CV3 prompts when interaction matters, noting kyutai/pocket-tts strength on GPU and CPU in the source. Then listen, vote with HF login where appropriate, and test on real scripts.
This sequence respects the leaderboard as a measurement design rather than a prize ceremony. The English leaders hexgrad/Kokoro-82M, Supertone/supertonic-3, and fishaudio/s2-pro mark strong intelligibility starting points Per HF, per the Hugging Face blog article. The multilingual names k2-fsa/OmniVoice, fishaudio/s2-pro, and FunAudioLLM/Fun-CosyVoice3-0.5B-2512 mark language breadth starting points Per HF, per the same source. Final choice still depends on licensing, voice roster, runtime budget, and listener approval in the target use case.
FAQ
What is the Open TTS Leaderboard?
It is an open TTS ranking that uses objective intelligibility, speed, and speaker similarity metrics to compare models repeatably, per the Hugging Face blog article Open TTS Leaderboard by Eric Bezzam, Steven Zheng, Eustache Le Bihan, and mrfakename dated Sep 30 2026.
How is the default English ranking computed?
The default ranking is a macro average WER on English Seed TTS Eval plus CV3 Eval in zero shot mode, per the Hugging Face blog article. Lower error rates point to clearer preservation of the intended words in those sets.
Which models lead the views named in the source?
English leaders are hexgrad/Kokoro-82M, Supertone/supertonic-3, and fishaudio/s2-pro Per HF, strong multilingual names include k2-fsa/OmniVoice, fishaudio/s2-pro, and FunAudioLLM/Fun-CosyVoice3-0.5B-2512 Per HF, and kyutai/pocket-tts is noted for streaming on GPU and CPU Per HF, per the Hugging Face blog article.
How are speed and speaker similarity measured?
Speed uses inverse RTF for batched offline synthesis on H200 plus time to first audio streaming with batch size 1 on H200 and CPU, while speaker similarity uses cosine SIM of WavLM embeddings, per the Hugging Face blog article.
Does a top rank mean the most natural voice?
No. The article states that the board does not replace human preference and that WER and SIM proxy intelligibility and identity, not naturalness, per the Hugging Face blog article. Use the Listen tab to compare outputs before choosing.
Where are the eval scripts published?
The eval scripts are open sourced under the repository name github.com/huggingface/open-tts-leaderboard in plain text here, per the Hugging Face blog article. That path is named without a link in this guide by rule.
Bottom line
The Open TTS Leaderboard gives open voice models a faster and more repeatable scoreboard, per the Hugging Face blog article Open TTS Leaderboard by Eric Bezzam, Steven Zheng, Eustache Le Bihan, and mrfakename dated Sep 30 2026. More than 8K TTS models on the Hub made fragmented evaluation a real bottleneck, while arena voting with Elo or Bradley Terry methods faced hosting, placement, and consistency limits. Objective WER and CER via Qwen3 ASR, inverse RTF batched offline on H200 plus time to first audio streaming batch 1 on H200 and CPU, and cosine SIM of WavLM embeddings compress evaluation from weeks to hours.
Start from the question you need answered. Use the default macro average WER on English Seed TTS Eval plus CV3 Eval zero shot when English clarity is the goal, with hexgrad/Kokoro-82M, Supertone/supertonic-3, and fishaudio/s2-pro as the leader group Per HF. Use multilingual toggles with CJK on CER when language breadth matters, with k2-fsa/OmniVoice, fishaudio/s2-pro, and FunAudioLLM/Fun-CosyVoice3-0.5B-2512 as strong names Per HF. Add the cloning toggle with SIM when identity matters, noting bosonai/higgs-tts-3-4b and openbmb/VoxCPM2 improved with reference audio Per HF. Check the Streaming tab TTFA median on 50 CV3 prompts with warm up dropped when interaction matters, noting kyutai/pocket-tts on GPU and CPU Per HF.
Then listen before you ship. The Listen tab compares outputs and collects votes with HF login, which adds the human preference signal that WER and SIM cannot provide. That final step catches flat delivery, awkward pacing, and voice mismatch that numbers miss. A shortlist from metrics plus a listening pass on real scripts will beat reliance on any single rank.
Sources
All model names, metric definitions, hardware references, evaluation sets, and tab behaviors in this guide come from the Hugging Face blog article Open TTS Leaderboard by Eric Bezzam, Steven Zheng, Eustache Le Bihan, and mrfakename dated Sep 30 2026. Hugging Face is named here in prose without hyperlinking except for the single Sources link required by house style. The eval code repository name github.com/huggingface/open-tts-leaderboard is given in plain text without hyperlinking by rule.