
A new study by USC researchers reveals that advanced audio AI models excel at transcribing spoken words but fail to understand tone, emotion, and other nonverbal cues that carry meaning in real conversation.
The team tested leading models and found they often ignore acoustic details like pitch and emotion, treating text as the dominant signal.
The researchers developed two novel techniques that dramatically improved the models' ability to interpret these paralinguistic cues, raising accuracy from 17% to 65%.
What happened
USC researchers discovered that audio large language models (LLMs)—AI systems that process spoken input—fail to interpret paralinguistic cues like tone, emotion, and pitch. The team tested NVIDIA's Audio Flamingo 3 and Alibaba's Qwen2-Audio and found both struggled when spoken words contradicted the emotional content of speech. The paper, "Do Audio LLMs Listen or Read? Analyzing and Mitigating Paralinguistic Failures with VoxParadox," was accepted to the International Conference on Machine Learning (ICML) 2026.
Why it matters
Audio LLMs treat language as primary and audio cues as secondary, often missing what is meant rather than what is said. When a speaker sounds happy but says "I am sad," the model incorrectly concludes sadness. This blind spot poses real risk in sensitive applications like health assessments or counseling, where misreading emotion could lead to harmful responses.
What to watch
The team developed two techniques—Prompt-Conditioned Layer Mixer (PCLM), which prioritizes relevant audio layers based on the question, and Direct Preference Optimization (DPO), a retraining method that aligns models with acoustic evidence. Together, the methods raised accuracy on paralinguistic tasks from 17% to 65%.
Led by USC professor Mohammad Soleymani, the research project began in August 2025 and produced the paper "Do Audio LLMs Listen or Read? Analyzing and Mitigating Paralinguistic Failures with VoxParadox," accepted to ICML 2026. The team included Soleymani's PhD students Ashutosh Chaubey and Jiacheng Pang. The work was motivated by a straightforward observation: AI can now hold voice conversations, receive spoken prompts, and analyze audio in real time through systems like ChatGPT and Gemini. Yet these audio LLMs—multimodal systems that convert speech into numerical vectors for processing—consistently fail to interpret meaning beyond the literal words. They are functionally "tone-deaf," missing the emotional, social, and contextual information that flows from how something is said.
To expose this blind spot, Soleymani's team designed stress tests using a benchmark called VoxParadox, which presented 10 different paralinguistic tasks: biometric challenges such as estimating age or gender and counting speakers, and prosodic and acoustic challenges such as identifying emotion, intonation, pitch, and volume. The critical twist was that spoken words intentionally contradicted the audio content. For instance, a speaker might sound cheerful but say "I am sad." When the model answered incorrectly—concluding the speaker was sad—it revealed that the model was prioritizing text over audio. Testing NVIDIA's Audio Flamingo 3 and Alibaba's Qwen2-Audio, both models struggled significantly, confirming a systemic problem.
Probing the models' internal layers revealed why. As audio data moved through the encoder toward language-processing components, acoustic details such as pitch and tone were progressively "stripped away"—a process the team called "representation degradation." By the final decision-making layer, much acoustic information had vanished. More troubling, even when correct acoustic information still existed in the model's internal representations, the decision-making process was so heavily biased toward language that it simply ignored the available audio evidence—a "utilization gap."
The solution came in two parts. First, the team developed Prompt-Conditioned Layer Mixer (PCLM), a module that examines the user's text prompt and determines which parts of the audio the model should "listen" to. Rather than relying only on the final encoder layer, PCLM draws information from multiple layers simultaneously—accessing early layers that capture basic acoustic details like pitch and volume, and deeper layers that understand vocabulary. For a question like "Is this speaker angry?" PCLM recognizes that emotional cues matter more than spoken words and draws more information from early acoustic layers. Soleymani described this as working "like a hearing aid" that "boosts the volume of the audio cues that are most relevant to your question." Second, the team applied Direct Preference Optimization (DPO), a post-training method that showed the model pairs of responses and taught it to prefer answers that better matched the audio, encouraging reliance on what it hears rather than textual shortcuts. Together, PCLM and DPO produced what the researchers called "very significant improvements." Accuracy on paralinguistic tasks jumped from 17% to 65%—a more than threefold gain. Soleymani noted that by identifying this new problem, the work opens the door for other researchers to recognize, study, and develop even better solutions, creating awareness of a blind spot the field had not fully understood before.
The study reveals a fundamental architectural imbalance in how today's audio AI systems process multimodal information. Audio LLMs are designed to handle both speech and text, but they are built on language models that inherently prioritize linguistic content over raw acoustic signals. As audio data flows through the model's neural layers toward the language-processing components, acoustic details are progressively stripped away—a phenomenon the researchers term "representation degradation." Even when the correct acoustic information exists in the model's internal representations, the decision-making process is so heavily biased toward language that it ignores the available audio evidence, a gap the team calls "utilization gap." This architectural weakness has consequences: in real-world scenarios where tone, emotion, and emphasis carry meaning—such as when someone says they are sad in an upbeat voice—the model confidently misinterprets intent, potentially leading to harmful outcomes in sensitive applications.
The two-part solution addresses both the loss of acoustic information and the model's reluctance to use it. PCLM acts as a dynamic router, examining the user's question and selectively drawing acoustic details from early processing layers (which preserve pitch and tone) rather than relying solely on deeper language layers. DPO retrains the model to align its reasoning with acoustic evidence, encouraging it to trust what it hears rather than textual shortcuts. Together, these techniques raised paralinguistic accuracy from 17% to 65%—a dramatic leap that suggests the models' failure was not inevitable but rather a design and training choice. The research opens a path for the field to recognize and address similar blind spots in how AI systems weight different modalities.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
The AI news that matters, in one minute each morning.
Sign up free