AIToday
Large Language ModelsRoboticsTHE DECODERPublished: Sep 13, 2026, 01:00 JST2 min read

GPT-6 Astra tops StationeryBench 7 to 0 over MolmoAct2

GPT-6 Astra tops StationeryBench 7 to 0 over MolmoAct2

3 Key Points

  1. What happened

    On StationeryBench, OpenAI's GPT-6 Astra fully completed 7 of 100 dual-arm robot tasks versus zero for Ai2's MolmoAct2, with median progress scores of 46 and 12 out of 100 across 200 trials.

  2. Why it matters

    Researcher Yoav Artzi calls Astra a "step change in spatial reasoning," and on the unpublished REMAP benchmark it reaches accuracy close to human level — though Artzi notes Astra still falls short of humans elsewhere.

  3. What to watch

    Artzi suspects OpenAI trained Astra on large amounts of 3D data such as Blender scenes, which would explain its gains on 3D tasks; the results, videos and code are on GitHub.

WHO IT HITSRobotics and embodied-AI teams evaluating dual-arm manipulation will likely treat StationeryBench as a new reference point for spatial reasoning, though the benchmark covers only five desk-object tasks.

Ask the AI about this article →

Summaries like this, in your inbox every morning.

Context & Analysis

StationeryBench is an early, small-scale test, but the gap it shows is stark: across 200 trials on the same dual-arm YAM robots, OpenAI's GPT-6 Astra finished 7 of 100 tasks while Ai2's MolmoAct2 finished none. Astra's median progress score of 46 out of 100, against MolmoAct2's 12, suggests the difference is not only about completing a task but about how far each model gets toward it.

The benchmark itself covers only five desk-object tasks — uncapping a marker, pouring out paper clips, passing a ruler between two robot arms — so it is a narrow slice of manipulation. Still, the results come with videos and code published on GitHub, and they follow OpenAI's stated long-term plans to build its own consumer robots, which makes this an early signal of how its models might perform outside text and images.

Outside researcher Yoav Artzi's description of a "step change in spatial reasoning" rests partly on the still-unpublished REMAP benchmark, where Astra reaches accuracy close to human level. Artzi's own caveat that "even ASTRA doesn't get to what humans do in other scenarios," and his suspicion that OpenAI trained on large amounts of 3D data such as Blender scenes, suggest the outcome hinges on how much of this gain is general versus tied to 3D-heavy training.

FAQ
What is StationeryBench?
It is a new robotics benchmark that tests models across five desk-object tasks, such as uncapping a marker or passing a ruler between two robot arms, using dual-arm YAM robots over 200 trials.
How did GPT-6 Astra and MolmoAct2 compare?
Astra fully completed 7 out of 100 tasks; MolmoAct2 completed zero. On median progress score, Astra reached 46 out of 100 versus MolmoAct2's 12.
Why does Astra seem better at 3D tasks?
Researcher Yoav Artzi suspects OpenAI trained the model on large amounts of 3D data such as Blender scenes.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Huang: 40,000 staff, 4M agents next at NvidiaYahoo Finance AI · 1h ago
  • Nvidia in talks to anchor Anthropic's $100 billion IPOYahoo Finance AI · 1h ago
  • KAIST study: AI reasoning steps separable inside modelsTHE DECODER · 1h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleMark Penn: Five fixes to save a billion and a half hours