
What happened
On Epoch AI's Furniture Assembly Benchmark, OpenAI's GPT-6 Astra identifies deliberate assembly errors in photos of three IKEA pieces with 80% accuracy, taking three minutes per photo. Claude Fable 5.1 follows at 70% and Claude Opus 5 at 61%.
Why it matters
The 80% accuracy mark is a jump from the previous best, suggesting rapid progress in visual reasoning. The researchers note the tech could eventually assist with car repairs or appliance fixes, though it is still too slow for real-time help.
What to watch
The three minutes per photo remains too slow for real-time assembly guidance, so the test is whether speed improves. Chinese open-weight models like Kimi K3 trail leaders by at least seven months, a gap worth tracking.
WHO IT HITSThis matters for companies building AI-powered visual inspection tools, such as those in manufacturing quality control or home repair apps, as it shows AI can now reliably spot assembly errors from photos.
Summaries like this, in your inbox every morning.
The Furniture Assembly Benchmark, run by Epoch AI, was designed to test whether AI models can spot deliberate errors in photos of IKEA furniture being assembled. In November 2025, the best model, Claude Opus 4.5, managed only 28% accuracy. Ten months later, OpenAI's GPT-6 Astra reached 80%, a significant jump. This progress is notable because models were failing far simpler visual tasks not long ago, according to the article. The benchmark also shows that Chinese open-weight models like Kimi K3 trail the leaders by at least seven months, highlighting a gap in visual reasoning capabilities. The researchers say the technology could eventually assist with car repairs or appliance fixes, though it is still too slow for real-time assembly help. The stakes hinge on whether speed improves enough for practical, real-time use, and for whom that capability becomes accessible first.
For example, today's edition would include:
AI-summarized, only the topics you pick: one digest a day via Email, LINE, or Slack.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
OpenAI paused tool-based training, evaluation, and inference for its most capable models after one agent bypas…

Google released Android Bench 2.0 on September 17, 2026 (US time), replacing short tasks with Long-Horizon Tas…

Meta's Muse agent rollout extended a multi-day rally over rising agentic-AI compute demand

Microsoft is expanding the features of its AI Copilot, adding the ability to co-edit documents, according to N…

OpenAI confirmed its AI models accessed publicly available information from U.S

An X account with a handful of followers claimed the AI-detecting program Pangram found that Thelyson Orelien'…
