Voice AI Is Failing
Is voice AI actually human-level yet? Hume's massive new study reveals why today's models still struggle to truly listen. 🎧
What the video says
We are told that voice AI is rapidly approaching human-level performance, yet anyone who actually talks to these models knows that something still feels deeply off. To bridge this gap, Hume has launched a landmark benchmark called Real World Voice EQ to measure the true human quality of voice AI.
While systems have become excellent at speaking, Hume argues they are still failing to truly listen, frequently missing critical paralinguistic cues like tone, hesitation, and emotional pacing. Traditional metrics increasingly overestimate how well these models perform in real-world conversations, where background noise and accents disrupt their flow.
To build a more meaningful standard, Hume collected over 1 million individual human ratings, including 785,000 text-to-speech assessments, This massive dataset exposes the limits of automated evaluators, revealing that speech language models struggle with subjective judgments, like whether a voice fits an acting role. In total, Real-World Voice EQ evaluates more than 40 leading models across 15 key dimensions to prove that technical speed is no longer enough.
Ultimately, Hume's data shows that no single system configuration ranked in the top 5 across all 8 major capability groups. Proving just how far voice AI has to go.