
What happened
On StationeryBench, OpenAI's GPT-6 Astra fully completed 7 of 100 dual-arm robot tasks versus zero for Ai2's MolmoAct2, with median progress scores of 46 and 12 out of 100 across 200 trials.
Why it matters
Researcher Yoav Artzi calls Astra a "step change in spatial reasoning," and on the unpublished REMAP benchmark it reaches accuracy close to human level — though Artzi notes Astra still falls short of humans elsewhere.
What to watch
Artzi suspects OpenAI trained Astra on large amounts of 3D data such as Blender scenes, which would explain its gains on 3D tasks; the results, videos and code are on GitHub.
WHO IT HITSRobotics and embodied-AI teams evaluating dual-arm manipulation will likely treat StationeryBench as a new reference point for spatial reasoning, though the benchmark covers only five desk-object tasks.
Ask the AI about this article →
Summaries like this, in your inbox every morning.
StationeryBench is an early, small-scale test, but the gap it shows is stark: across 200 trials on the same dual-arm YAM robots, OpenAI's GPT-6 Astra finished 7 of 100 tasks while Ai2's MolmoAct2 finished none. Astra's median progress score of 46 out of 100, against MolmoAct2's 12, suggests the difference is not only about completing a task but about how far each model gets toward it.
The benchmark itself covers only five desk-object tasks — uncapping a marker, pouring out paper clips, passing a ruler between two robot arms — so it is a narrow slice of manipulation. Still, the results come with videos and code published on GitHub, and they follow OpenAI's stated long-term plans to build its own consumer robots, which makes this an early signal of how its models might perform outside text and images.
Outside researcher Yoav Artzi's description of a "step change in spatial reasoning" rests partly on the still-unpublished REMAP benchmark, where Astra reaches accuracy close to human level. Artzi's own caveat that "even ASTRA doesn't get to what humans do in other scenarios," and his suspicion that OpenAI trained on large amounts of 3D data such as Blender scenes, suggest the outcome hinges on how much of this gain is general versus tied to 3D-heavy training.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
On Nvidia's latest earnings call, CEO Jensen Huang said AI crossed an inflection point last month, with most A…

Nvidia is reportedly discussing anchoring Anthropic's planned $100 billion IPO at a valuation near $2 trillion

A KAIST and Naver AI Lab study found that reasoning operations like extraction, decomposition, formula recall…

Reuters reports Nvidia is in talks to invest up to $10 billion in Anthropic's planned IPO as an anchor investo…

OpenAI's Eric Provencher recommends reviewing skills, AGENTS.md, and task prompts when switching to GPT-6 Astr…

Anthropic CEO Dario Amodei published an essay proposing a three-step plan to slow frontier AI progress, and sa…
