AIToday
Large Language ModelsTHE DECODERPublished: Aug 22, 2026, 19:01 JST3 min read

World models that ignore beliefs predict wrong actions, research finds

World models that ignore beliefs predict wrong actions, research finds

Key takeaway

  • Researchers show that AI world models predicting human actions must track mental states like beliefs and emotions, not just physical objects.

  • Testing eight language models on 448 decision scenes, they found that a framework called MWM improved prediction accuracy to 87.9 F1 score from 63.3 with direct answers alone.

  • The bottleneck is simulating how mental and physical states change together.

3 Key Points

  1. What happened

    Researchers developed Mental World Modeling (MWM), a framework that extends AI world models by tracking mental variables like beliefs, attention, goals, and emotions alongside physical states. They built MENTIS, a training-free implementation, and tested it on Menti-Bench, a dataset of 448 decision scenes with responses from eight language models including GPT-5.6-Sol, GPT-4.1, Claude Opus 4.8, and Claude Haiku 4.5.

  2. Why it matters

    Current AI models fail to predict human actions in scenarios where hidden mental states matter—such as when a person searches for a cup moved without their knowledge. The MWM framework shows this gap is not simply fixed by sampling more answers: the weakest model with MWM (GPT-4.1, F1 score 84.9) beats the strongest model using direct answers with self-consistency (GPT-5.6-Sol, 83.6). In interpersonal scenes, MWM improves predictions by 26.4 F1 points; in object-focused scenes, by 14.0 points.

  3. Why it matters

    The research reveals that roughly 80 percent of the remaining gap to human performance (humans scored 98.5) traces to errors in simulating state transitions—how the coupled physical-mental world changes. Removing the mental channel alone costs an average of 12.1 F1 points; removing the physical channel costs 16.5 points. The authors argue future improvements should focus on better simulation of these transitions, not raw model capacity.

Ask the AI about this article →

Context & Analysis

World models—systems that predict what happens next in a scene—have become a major research focus after the success of large language models. Demis Hassabis, who recently stepped down as operational head of Google Deepmind, has said he spends most of his research time on world models and expects a 'ChatGPT moment' for the field. However, the field still lacks agreement on what counts as a world model: proposals exclude text-to-video models like Sora for lacking feedback loops with reality, while others criticize the generative approach entirely in favor of abstract representations.

The MWM framework addresses a specific failure mode that prior research had already identified: language models struggle with Theory of Mind tests (inferring what others believe) and perform even worse at tracking world states. A Meta FAIR lab team and researchers from Washington and Carnegie Mellon had shown these gaps directly. Interestingly, Anthropic's recent discovery of an internal scratchpad in Claude—a hidden layer of word-like thoughts that support multi-step reasoning—suggests that something analogous to mental state tracking may be forming inside models themselves, even if not explicitly engineered.

FAQ

What exactly is Mental World Modeling and how does it work?
MWM extends AI world models by tracking mental variables—beliefs, attention, goals, intentions, emotions, norms, and social relationships—alongside physical objects and states. Every action is split into a physical carrier (speaking, pointing, grasping) and a mental payload (comforting, deceiving, rejecting), allowing the system to understand that the same gesture can mean different things depending on intent.
How much better is MWM than standard language model responses?
On the Menti-Bench dataset, direct answers from language models scored 63.3 F1. Self-consistency (asking six times and picking the most common answer) improved it to 77.9. The full MWM pipeline reached 87.9, while humans scored 98.5 under the same protocol.
Where do AI systems still struggle with this framework?
About 80 percent of the remaining gap to human performance comes from prediction errors in state transitions—simulating how both mental and physical states change when an action occurs. Providing correct state transitions alone would add 3.5 F1 points, the single largest gain possible from fixing any individual pipeline stage.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next article80-year-old AC maker Vertiv surges to $109B on AI cooling dominance