
What happened
LEGO-Anything has a coding agent write executable Blender code from a single photo, and in the accompanying LEGO-Bench the top model, GPT-6 Astra, hit 53.4 percent accuracy on indoor scenes but just 39.6 percent outdoors.
Why it matters
Even the strongest tested model's scenes fall short of a faithful reconstruction, so the team concludes refinement should rely on concrete measurements rather than the agent's own judgment.
What to watch
The gap between a working scene and a faithful one hinges on whether agents can stop undoing their own progress, since the researchers found revisions that reversed earlier gains and unreliable self-assessment; the article gives no timeline for closing it.
WHO IT HITSTeams using AI coding agents for 3D content — game and architectural visualization pipelines that need editable Blender scenes — get a first benchmark showing usable output but limited geometric fidelity.
Summaries like this, in your inbox every morning.
The work is built around a simple shift: instead of generating a 3D scene directly, LEGO-Anything has a coding agent produce a Blender program. Because the output is code, objects, geometry, layout, and camera position are captured explicitly, and the scene can be run, checked, and modified. To measure how well this actually works, the researchers built LEGO-Bench from professionally built simulator scenes, giving the inputs a natural look while keeping exact geometry, depth, and object assignments as an automated answer key.
The benchmark exposed a sharp split. All six tested GPT configurations delivered a working scene almost every time, but accuracy varied widely: GPT-6 Astra, the best tested model, hit 53.4 percent on indoor scenes and 39.6 percent on outdoor scenes, while weaker configurations scored around 15 percent. More complexity meant lower accuracy, and outdoor scenes proved harder than interiors. When the researchers increased the models' reasoning budget, the GPT-6 variants improved a lot — Astra jumped from 32.3 to 61.8 percent on an office test subset.
Digging into why, the team found poor initial attempts, revisions that undid earlier progress, and unreliable self-assessment as the most common issues. That last point is the most telling: when models had to pick which of two versions better matched the original, their judgments landed near or below chance, even for their own scenes. The plugin they built in response improved all six models, with weaker agents gaining up to 62.7 percent while the top model gained only about two percentage points. The open question is whether these scenes can serve as a basis for standard vision tasks; the authors report usable but unremarkable results, with a big gap remaining between a working result and a faithful reconstruction.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Manulife Financial Corporation launched a first-of-its-kind CoverMe travel insurance plugin in ChatGPT in late…

Google said free Gemini app users will be restricted to the "Flash-Lite" model starting October 9; "Flash" nee…

Indie developer Robert Varadan argues that AI models like Opus 5.5 and GPT-6 Astra can clone game demos from a…

A developer published Rai, a small Rust engine that runs language models on an ordinary PC's CPU, using only t…

The developer published fk2000/minecraft-ai-bot, which uses Jev to assemble instructions, Gemini to pick from…

On September 22, 2026, Anthropic announced Claude Opus 5.5 at $4 input / $20 output per million tokens, then O…
