AIToday
Semafor TechPublished: Apr 9, 2026, 04:00 JST1 min read

New ARC-AGI-3 benchmark test reveals that even the most advanced AI models fail to solve game-like puzzles, scoring below 1% on tasks designed to detect when artificial general intelligence arrives.

New ARC-AGI-3 benchmark test reveals that even the most advanced AI models fail to solve game-like puzzles, scoring below 1% on tasks designed to detect when artificial general intelligence arrives.

3 Key Points

  1. ARC-AGI-3 test uses game-like puzzles that require AI systems to solve problems on the fly without prior training

  2. Current top-performing AI models score below 1% on the benchmark, indicating significant gap from human-level general intelligence

  3. Test was created by a research foundation as a potential early warning system to detect when AGI has been achieved

  4. The benchmark focuses on reasoning and problem-solving abilities that distinguish general intelligence from narrow, specialized AI capabilities

Ask the AI about this article →

Get AI news like this every morning

For example, today's edition would include:

  • AI agents outpace human security teams, forcing endpoint defensesSiliconANGLE AI · 1h ago
  • PlayNitride eyes Micro LED optical comms in 2 yearsDIGITIMES Asia · 1h ago
  • OpenAI report misses cultural failures behind AI hackMITテクノロジーレビュー · 1h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Next articleStability AI enters commercial market with Brand Studio, enabling businesses to generate on-brand AI images at scale