AIToday
Large Language ModelsAI Safety & AlignmentMIT Technology Review AIPublished: Aug 26, 2026, 19:00 JST1 min read

AI puzzles reveal blind spots, human wins

AI puzzles reveal blind spots, human wins

Key takeaway

  • AI models excel at many puzzles but fail at spatial reasoning and visual tests.

  • In late 2024, best models solved only 18% of Connections puzzles, yet improved by 2025.

  • Humans still beat them in certain areas.

3 Key Points

  1. What happened

    AI models improve rapidly at puzzles, yet still fail at spatial reasoning and visual puzzles. In late 2024, the best models solved only 18% of New York Times Connections puzzles; by early 2025, some solved them near perfectly.

  2. Why it matters

    These tests expose where machine and human cognition differ. Subtle changes to classic riddles often trip models up, and visual puzzles remain a weak spot, showing limits despite advances.

  3. What to watch

    Try the seven puzzles yourself, from mental rotation to logic grids. Some are tricky for humans; others highlight AI's surprising failures, such as on SimpleBench where top-tier models trip.

Ask the AI about this article →

Context & Analysis

Puzzles have been central to AI development since the 1959 checkers-playing algorithm by IBM's Arthur Samuel, and now serve to reveal model limitations. While AI solves Connections puzzles near perfectly by early 2025, significant gaps remain.

Spatial reasoning is a key weak spot; language models cannot manipulate 3D objects like architects or mechanical engineers. Similarly, visual puzzles like ARC-AGI show models often use non-generalizable rules, unlike humans who draw on simple visual concepts.

Scale also matters: Apple researchers found models master simple Tower of Hanoi and river-crossing puzzles but falter at six or more disks or people. Logic grid puzzles from the University of Washington, Stanford, and Allen Institute show similar struggles, though commentators question whether this reflects a unique reasoning limitation or normal error as complexity increases.

FAQ

What is ARC-AGI and why is it important?
ARC-AGI is a famous puzzle-based benchmark requiring models to infer abstract rules from examples. Models improve when grids are given as numbers, yet some puzzles still stump them.
Why do models fail at Knights and Knaves puzzles?
A 2024 study from Google and the University of Illinois showed models struggle when puzzles closely resemble training data. They may whiz by key differences and respond with what they memorized.
MIT Technology Review AIRead Original Article

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleApple takes AI compute out of data centers with new Macs