← Back to list

Every TTS Model Fails the Same Way. Humans Caught It. Benchmarks Didn’t.

Published by VoiceArena | Josh Talks AI

Vishaljha · 2026-05-30 14:01 · 46 claps · 7.4 min read
#artificial-intelligence #machine-learning #technology #speech-technology #voice-ai
Open on Medium ↗
Wiki topics: EVAL · Evaluation & Benchmarks MM · Multimodal & Generative Media ML · Machine Learning AI · AI · General EDU · Education & Learning

Every TTS Model Fails the Same Way. Humans Caught It. Benchmarks Didn’t.

Published by VoiceArena | Josh Talks AI

Across 44,387 human votes, six languages, and seven of the most capable text-to-speech models available today, one failure pattern appeared consistently — in English, Hindi, Arabic, Japanese, Portuguese, and Vietnamese.

Every model. Every language. Same failure.

None of the standard automated benchmarks flagged it.

The analysis includes evaluations from six leading text-to-speech (TTS) systems that were tested across all six languages in the VoiceArena dataset: Google Gemini 3.1 Flash TTS, ElevenLabs Eleven v3, OpenAI GPT-4o Mini TTS, Microsoft Azure Dragon HD Omni, Cartesia Sonic-3, and xAI Grok TTS.

The full VoiceArena leaderboard is updated continuously as new evaluations are collected. Readers interested in current rankings and Elo scores can explore the live leaderboard here.

[View the VoiceArena Leaderboard →]

What We Measured

VoiceArena uses pairwise human evaluation to rank TTS models. Two models receive the same input sentence. Both generate audio. A human listener — who does not know which model produced which audio — listens to both and votes for the better one, often leaving a written comment explaining why.

This is the same structure as Chatbot Arena for LLMs, applied to voice. The key difference from automated benchmarks is that the human is evaluating the full listening experience — not a waveform, not a phoneme alignment score, not a word error rate. Just: does this sound like something a real person would say?

This evaluation ran across six languages:

The vote log includes not just win/loss outcomes but written human comments — what the evaluator actually heard, what bothered them, and why they chose one model over another. This comment layer is where the most useful signal lives.

The Finding: Filler Words Break Every Model

VoiceArena evaluates conversational speech, not just clean declarative sentences. This means the test corpus includes sentences with filler words — “umm,” “uh,” “err,” and their equivalents in other languages.

These are the sounds humans make when they are thinking mid-sentence. They appear constantly in customer service calls, voice assistants, real-time agents, and any TTS use case that mimics natural conversation.

Every model tested failed on filler words. Not occasionally. Consistently. Across all six languages.

In English, human evaluators wrote:

“The word ‘uh’ sounds unnatural.” “The ‘umm’ at 0:08 is short and makes it sound robotic.” “There was pausing between words that were unnatural, which makes the voice sound robotic.” “When the audio makes the ‘uh’ sound, it is drawn out and becomes buzzy, making it sound unnatural.”

In Hindi, the same pattern appeared in a different form. Evaluators wrote in both English and Hindi:

“Filler word ko sahi pronounce nhi kiya jisse robotic sound karrha hai.” (The filler words are not pronounced correctly, making it sound robotic.) “The pronunciation and content are correct, but the fillers like ‘umm/uh’ sound slightly unnatural, affecting the natural flow.”

In Arabic, where the filler words are أم and أه, the failure was flagged against every single model in the evaluation — Microsoft Azure Dragon HD Omni, Google DeepMind Gemini 3.1 Flash TTS, OpenAI gpt-4o-mini-tts, Cartesia Sonic-3, and ElevenLabs Eleven v3. Comments like “unnatural tone أم و أه” appear dozens of times in the Arabic vote log. Evaluators noted not just that the words sounded wrong, but that the intonation of the filler did not match the surrounding sentence.

In Japanese, evaluators flagged the filler word えーと (e-to) for unnatural intonation. One evaluator wrote:

“It makes minor mispronunciation mistake ‘ついてきなす’. And the intonation of ‘えーと’ sounds unnatural.”

This is not a random distribution of individual model failures. It is a structural pattern: no current TTS model has solved how to produce filler words that sound natural inside running speech.

Why Automated Benchmarks Miss This

Standard TTS benchmarks — MOS scores, word error rates, character error rates, speaker similarity scores — measure production quality on clean reference sentences. They do not evaluate whether a model can produce a convincing “umm” mid-sentence in a way that sounds like a real person who is actually thinking.

Filler words are not in most TTS training prompts. They are not in most evaluation scripts. When they do appear, automated tools grade them on pronunciation accuracy, not on whether the pause length, pitch contour, and trailing vowel match how a real speaker would produce that sound in context.

Human evaluators catch this immediately. The failure is perceptual — it breaks the illusion of naturalness. A voice can pronounce every word correctly and still sound mechanical if the filler words are wrong.

VoiceArena’s methodology, which uses real human evaluators on conversational sentences, is one of the few evaluation setups where this failure can surface. It would not appear in a benchmark built entirely on clean read speech.

The Performance Data

Filler word failure is the shared failure. But models differ significantly in how they fail elsewhere.

Win rates by model and language (excluding ties):

Win rate = percentage of non-tie matchups won. Ties are excluded from this calculation.

Google leads every language. The gap is widest in Japanese (63%) and Arabic (61.5%), and narrowest in English (40.8%). ElevenLabs consistently holds second position across languages. The gap between first and second is smaller in English than in any non-English language — suggesting English is a more competitive benchmark environment where multiple models have invested significantly.

Model-Specific Failure Patterns

Beyond filler words, each model has a distinct failure signature. The human comments allow these to be identified with specificity.

Cartesia Sonic-3 — Speed and Rhythm

Cartesia is the only model where speed consistently appears as a primary complaint. Human evaluators across English and Japanese flagged it repeatedly:

“This is way too fast and does not allow for emphases within sentences.” “The pace is fast. Voice is robotic.” “Too fast through the entire thing. Makes it sound very robotic.” “The speech is too fast and feels rushed.”

In the English vote log alone, speed-related comments appear 321 times for Cartesia — more than for any other model on any other dimension. Cartesia’s English win rate of 13.5% is the lowest of any model in any language except Microsoft’s Japanese performance.

Microsoft Azure Dragon HD Omni — Mispronunciation

Microsoft receives the highest mispronunciation complaint volume across the full dataset. Human evaluators identified specific words that were consistently mispronounced:

In English: tracking IDs, service codes (CHI, COMP, HVAC codes), proper names (Caoimhe, Saoirse, Cuyahoga), and dollar amounts.

“Mispronounced HVAC-CHI-4351. For the letter CHI, the model makes it the word ‘chi’ instead of saying the letters individually.” “Mispronounced the tracking ID 1Z999AA10123456784. All the characters were read incorrectly.”

In Japanese, Microsoft’s performance collapsed. Multiple evaluators wrote “several words are badly mispronounced, making the audio difficult to understand” — a phrase that appears in the Japanese vote log over forty times. Microsoft’s Japanese win rate is 8.0%, the lowest single figure in the entire dataset.

OpenAI gpt-4o-mini-tts — Robotic Tone

OpenAI receives the highest volume of robotic-voice complaints: 375 mentions across the evaluated languages. In Hindi, the phrase “unnatural robotic voice hai” (this has an unnatural robotic voice) appears in the vote log in very close variations dozens of times — suggesting a consistent perceptual quality that evaluators noticed and described with similar language.

“There was pausing between words that were unnatural, which makes the voice sound robotic.” “The audio is clear but flat and monotone sounding, making it unnatural.” “The ‘s’ sounds are drawn out and there is a little whistling when it is said.”

What This Means for Teams Building These Models

Three specific things follow from this data.

1. Add a filler word evaluation suite to your internal benchmarks.

If your model is being deployed in any conversational context — customer service, voice agents, real-time assistants — your current automated benchmarks are not telling you how your filler words sound. Build an evaluation set of 100–200 sentences that include mid-sentence fillers in every language you support. Run human listeners on it. The gap between your WER and your filler word perceptual score is likely wider than you think.

2. Test alphanumeric strings, codes, and proper names as a separate category.

The Microsoft and Google Arabic mispronunciation failures cluster around specific sentence types: tracking IDs, booking references, chemical formulas, product codes, personal names. These are extremely common in real deployment environments. A model that reads clean prose well but mispronounces “APL-RPR-DXB-2991X” is not production-ready for customer service. VoiceArena’s sentence design specifically includes these categories. Your internal evaluations should too.

3. Non-English performance requires non-English evaluators.

Google’s 63% Japanese win rate and 61.5% Arabic win rate did not come from English-language optimization. The gap between Google and the second-ranked model is substantially larger in non-English languages than in English. Teams building for global deployment cannot evaluate non-English performance with English-trained automatic metrics — and this data shows the performance difference is large enough to matter significantly.

Limitations of This Evaluation

This evaluation has constraints worth naming directly.

The evaluation uses one assigned voice per model per language. Models like ElevenLabs that offer many voices may perform differently across their voice catalog. This methodology evaluates a model’s default or assigned voice, not its ceiling.

Human evaluators vary in their background and linguistic competence. In some languages, the evaluator pool may not fully represent the range of regional accents and usage patterns that a deployed model would encounter.

Tie votes — which account for 36–48% of outcomes depending on language — indicate cases where evaluators found no meaningful difference between models. A high tie rate does not necessarily mean models are equal on all dimensions; it may mean that for certain sentence types, the differences are too small for untrained listeners to consistently detect.

Win rate alone does not capture the magnitude of wins. A model that wins by a small margin is counted the same as one that wins decisively. The comment data is a better signal for magnitude.

Closing Note

The filler word finding is not a minor quality issue. It is a marker of the gap between text-to-speech and speech synthesis. TTS systems are trained to convert text to audio. Real conversational speech includes sounds that are not text — pauses that carry meaning, fillers that signal cognitive load, hesitations that communicate uncertainty.

No model in this evaluation has solved that gap. The human evaluators noticed it in every language.

The question for teams building these models is not whether this failure exists. The vote log shows clearly that it does. The question is whether your current evaluation infrastructure would have told you.

This analysis is based on 44,387 pairwise human votes collected by VoiceArena across English, Hindi, Arabic, Japanese, Portuguese, and Vietnamese. Vote log data was collected between April and May 2026. VoiceArena uses ELO-based scoring with blind pairwise evaluation. Methodology details at voicearena.com/tts-methodology.

VoiceArena is built by Josh Talks AI, which develops voice AI infrastructure, datasets, and evaluations for model builders and AI labs.


메타데이터
post_id
45c3fbfa1ac7
slug
every-tts-model-fails-the-same-way-humans-caught-it-benchmarks-didnt-45c3fbfa1ac7
url
https://medium.com/@vishaljha154/every-tts-model-fails-the-same-way-humans-caught-it-benchmarks-didnt-45c3fbfa1ac7
canonical_url
https://medium.com/@vishaljha154/every-tts-model-fails-the-same-way-humans-caught-it-benchmarks-didnt-45c3fbfa1ac7
author_url
https://medium.com/@vishaljha154
status
ok
fetched_at
2026-06-09 15:37:30