
AI models excel at many puzzles but fail at spatial reasoning and visual tests.
In late 2024, best models solved only 18% of Connections puzzles, yet improved by 2025.
Humans still beat them in certain areas.
What happened
AI models improve rapidly at puzzles, yet still fail at spatial reasoning and visual puzzles. In late 2024, the best models solved only 18% of New York Times Connections puzzles; by early 2025, some solved them near perfectly.
Why it matters
These tests expose where machine and human cognition differ. Subtle changes to classic riddles often trip models up, and visual puzzles remain a weak spot, showing limits despite advances.
What to watch
Try the seven puzzles yourself, from mental rotation to logic grids. Some are tricky for humans; others highlight AI's surprising failures, such as on SimpleBench where top-tier models trip.
Ask the AI about this article →
Puzzles have been central to AI development since the 1959 checkers-playing algorithm by IBM's Arthur Samuel, and now serve to reveal model limitations. While AI solves Connections puzzles near perfectly by early 2025, significant gaps remain.
Spatial reasoning is a key weak spot; language models cannot manipulate 3D objects like architects or mechanical engineers. Similarly, visual puzzles like ARC-AGI show models often use non-generalizable rules, unlike humans who draw on simple visual concepts.
Scale also matters: Apple researchers found models master simple Tower of Hanoi and river-crossing puzzles but falter at six or more disks or people. Logic grid puzzles from the University of Washington, Stanford, and Allen Institute show similar struggles, though commentators question whether this reflects a unique reasoning limitation or normal error as complexity increases.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
Meta halted its planned AI-driven layoffs hours before the first round, after internal documents revealed a pl…

Sentence Transformers v6.0 adds a fourth model type, MultiVectorEncoder, for ColBERT-style late-interaction re…

Particle, the AI newsreader startup founded by former Twitter engineers, introduced Radar, a podcast search en…

Z.ai, the maker of the GLM series, confirmed that Ox Alpha, a model launched anonymously on OpenRouter, is the…

Bill Gates posted a long essay on Gates Notes proposing a “robot tax” and setting aside certain jobs as “Human…

Glean Technologies Inc. today unveiled Glean Tau, a desktop workspace connecting its enterprise AI to local fi…