The Illusion of “Naturalness” in Speech Evaluation

Layered speech waveforms with one curve highlighted in orange

Is “naturalness” the right metric for Text-to-Speech? For AI assistants, yes. But if you want an AI that can act: one that can panic, hesitate, or whisper a secret, the standard naturalness Mean Opinion Score (MOS) and commonly used metrics fall apart.

In our Interspeech 2026 paper, “Is Natural Always Appropriate?”[1], in collaboration with KTH, Imperial, and TUM München, we evaluate state-of-the-art text-to-speech (TTS) models across different “personas” to challenge a core assumption: that a more “human-sounding” voice is always the right fit for the job.

The Acting Problem

As a team of AI researchers and game-makers, we have been exploring and evaluating voice models for characters. Rather than a narrator or an AI assistant, we want AI voices that feel alive and make you believe in their universe and story.

Speech is fundamentally a one-to-many mapping. A single line like “Are you serious?” changes entirely based on context, pacing, and breath. It can be a genuine question, a dark threat, or a deadpan punchline. However, we have observed that current TTS models often fail once they are asked to act, rather than to read or to hold a single fixed speaking style. Acting requires range, flexibility, and contextual realisation, and that is precisely what these systems lack.

The problem is that standard evaluation metrics couldn’t tell us the difference between a good and a bad performance. We quickly learned that “naturalness” ended up being a red herring: a model could generate a flawlessly natural-sounding human voice that still completely misses the emotional truth of the scene. We decided to investigate this shift from naturalness to appropriateness, and identify which standard evaluation metrics fail and how under this new light.

The Human-Likeness Paradox

We conducted a systematic cross-domain perceptual analysis. We tested five voices from SOTA TTS models (Gemini-TTS, ElevenLabs, Kyutai-TTS, GPT4o-mini and Kokoro) across 150 human listeners, mapping their performance across five distinct personas.[†]

We found that a model’s ability to sound “human” does not predict its ability to sound appropriate for a given scenario.

Human-likeness × appropriateness, by persona

Spearman ρ between a sample’s human-likeness rating and its appropriateness rating for that persona.

Through our study, we identified four key findings:

Below, we show the profile of the different TTS and voices tested, which general preference benchmarks or human-likeness scores alone typically obscure. Every synthetic voice leans toward specific styles, and only the human recordings manage to capture the full range of expression. Note that these findings reflect a single evaluated voice per system rather than the theoretical limits of the underlying models.

Voice profiles — appropriateness

Systems shown

Mean appropriateness (1–5) as rated by 150 listeners. Sentences matched to each persona.

† The voice used in each system: Kokoro af_heart, GPT-4o-mini-TTS coral, ElevenLabs Bella (model eleven_multilingual_v2), Gemini-TTS Despina (gemini-2.5-flash-preview-tts), and Kyutai-TTS, prompted with speaker p037 from the EARS corpus.

The Metrics Blind Spot

The fields of TTS and TTS evaluation have notoriously struggled with evaluation. Hume AI recently released their Real World VoiceEQ benchmark[2], rightfully pointing out that legacy metrics like WER[3], PESQ[4], and DNSMOS[5] fail because they ignore emotion, tone, and paralinguistics.

We fully agree that we need to move beyond clinical metrics. But while evaluating “human expressivity” is a step forward, it still falls into a subtle trap: it assumes there is a universal standard for a “good” voice or performance.

As our paper shows, this assumption misses the bigger picture. A highly expressive, emotionally nuanced delivery will actively frustrate a user if they are talking to an AI assistant or a call center agent: appropriateness is entirely dependent on persona. Like how actors are casted for specific roles, you cannot measure a digital actor without knowing the role they are cast to play.

We tested these hypotheses and correlated standard automatic metric against human appropriateness ratings across the different downstream personas.

Automatic metrics × appropriateness, sentence level

Spearman ρ between automated metrics and human listeners. Orange = the metric agrees with listeners, blue = it disagrees. * marks significance. Rows marked ↓ are lower-is-better.

Current automated metrics and benchmarks fail to measure expressive speech:

Hear the Difference

To better understand what we mean, we suggest you the reader have a try yourself.

Listen to the examples below from our study. Notice how the traditional automated metrics (like UTMOS for audio quality) often reward the out-of-character delivery and penalize the appropriate acting. A voice can sound excellent in a vacuum and fall into the uncanny valley when it’s put against the wrong persona.

Actor

“Nobody respects the bucket!”

Listeners preferred

ElevenLabs

  • Appropriateness 3.24
  • Human-likeness 3.32
  • UTMOS 2.23
More human-like

Kyutai-TTS

  • Appropriateness 3.12
  • Human-likeness 4.32
  • UTMOS 2.61
Spontaneous speaker

“I… I mean, they’re just disgusting to me.”

Listeners preferred

ElevenLabs

  • Appropriateness 3.12
  • Human-likeness 3.52
  • UTMOS 3.42
More human-like

Gemini-TTS

  • Appropriateness 2.28
  • Human-likeness 3.68
  • UTMOS 3.37
AI assistant

“Your meeting starts in 30 minutes.”

Listeners preferred

Kokoro

  • Appropriateness 4.40
  • Human-likeness 2.72
  • UTMOS 3.96
More human-like

Kyutai-TTS

  • Appropriateness 2.72
  • Human-likeness 3.80
  • UTMOS 2.76

Beyond the Vacuum: Building Context-Aware Digital Actors

TTS isn’t a one-size-fits-all problem. Generating clean audio in a vacuum is no longer the benchmark. Real speech is driven by intent, context, and narrative stakes.

To capture this, synthesis must move beyond the traditional text-in, audio-out pipeline. Next-generation TTS is entirely context-driven. It can be controlled either through explicit emotional markup or a unified model that digests the full scope of a scene and conversation, deliberately generating paralinguistics like breaths and hesitations to shape the delivery.

At Iconic, we are tackling this challenge by bridging the gap between artificial intelligence and artistic performance. Because a standard pipeline optimizes the life out of a character, our researchers work directly alongside creative directors, writers, and voice actors to map personas deliberately. Rather than training generalist models, we build unified persona models in which the LLM’s understanding of a character’s psychological state and the voice engine’s acoustic delivery function as a single system.

The next great medium of interactive entertainment won’t be built with voices that simply sound human in a laboratory setting. It will be built with digital actors that know how to perform and act in context.

Learn More

If you want to learn more or chat, we will be presenting the paper at Interspeech 2026 in Sydney in September. You can check out the paper here!

In the meantime, you can check out our blog post on the Pressure Point tech demo to experience firsthand how we are applying these concepts to push the boundaries of interactive entertainment!

References

  1. D. Woszczyk, A. Triantafyllopoulos, J. Miniota, É. Székely, B. Schuller. “Is Natural Always Appropriate? Investigating Naturalness and Appropriateness Across Different Domains for TTS Evaluation.” arXiv:2606.31729, 2026.
  2. Hume AI. “Introducing Real World VoiceEQ: Measuring the Human Quality of Voice AI.”
  3. NVIDIA. “nvidia/parakeet-tdt-0.6b-v2.” Hugging Face, 2025. Used for the WER transcriptions.
  4. A. W. Rix et al. “Perceptual Evaluation of Speech Quality (PESQ) — A New Method for Speech Quality Assessment of Telephone Networks and Codecs.” ICASSP, vol. 2, 2001, pp. 749–752.
  5. C. K. A. Reddy et al. “DNSMOS: A Non-Intrusive Perceptual Objective Speech Quality Metric to Evaluate Noise Suppressors.” ICASSP, 2021, pp. 6493–6497.
  6. K. Baba et al. “The T05 System for the VoiceMOS Challenge 2024: Transfer Learning from Deep Image Classifier to Naturalness MOS Prediction of High-Quality Synthetic Speech.” IEEE Spoken Language Technology Workshop (SLT), 2024, pp. 818–824.
  7. R. Kubichek. “Mel-Cepstral Distance Measure for Objective Speech Quality Assessment.” Proceedings of IEEE Pacific Rim Conference on Communications, Computers and Signal Processing, vol. 1, 1993, pp. 125–128.
  8. C. H. Taal et al. “An Algorithm for Intelligibility Prediction of Time–Frequency Weighted Noisy Speech.” IEEE Transactions on Audio, Speech, and Language Processing, vol. 19, 2011, pp. 2125–2136.
  9. A. Vyas et al. “Audiobox: Unified Audio Generation with Natural Language Prompts.” arXiv:2312.15821, 2023.
  10. L. Barrault et al. “Seamless: Multilingual Expressive and Streaming Speech Translation.” arXiv:2312.05187, 2023. Source of the AutoPCP prosody-consistency measure.
  11. A. Kumar et al. “TorchAudio-Squim: Reference-less Speech Quality and Intelligibility Measures in TorchAudio.” ICASSP, 2023, pp. 1–5.
  12. S. Chen et al. “WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing.” IEEE Journal of Selected Topics in Signal Processing, vol. 16, 2022, pp. 1505–1518.
  13. Y. Yang et al. “Measuring Prosody Diversity in Zero-Shot TTS: A New Metric, Benchmark, and Exploration.” arXiv:2509.19928, 2025.
  14. K. Kilgour, M. Zuluaga, D. Roblek, M. Sharifi. “Fréchet Audio Distance: A Reference-Free Metric for Evaluating Music Enhancement Algorithms.” Interspeech, 2019.
  15. C. Minixhofer et al. “TTSDS2: Robust Objective Evaluation for Human-Quality Synthetic Speech.” The 13th Speech Synthesis Workshop, 2025, pp. 68–75.