
What happened
Researchers at MIT CSAIL and Toyota Research Institute developed SceneSmith, a system that uses three AI agents powered by GPT-5.2 to generate detailed 3D indoor environments. The agents—a designer, critic, and orchestrator—work together to create realistic scenes with up to six times more objects per space than prior methods, which robots can then practice tasks in before real-world deployment.
Why it matters
Creating training data for robots is labor-intensive and time-consuming. By generating rich, physics-accurate virtual environments from text prompts alone, SceneSmith lets engineers test robot behavior in simulation before deployment, saving real-world trial-and-error. In evaluation, a robot policy trained on real-world data successfully completed tasks in SceneSmith scenes—suggesting the simulations are realistic enough to transfer learning.
What to watch
The system currently takes multiple hours to generate a single scene due to the agents' detailed scrutiny of each object. Researchers note that increased computing power could dramatically improve efficiency, and they plan to expand to deformable objects like sponges if 3D libraries become available. The team presented findings as a spotlight at the 2026 International Conference on Machine Learning.
Summaries like this, in your inbox every morning.
The core challenge roboticists face is that robots learn best through experience, but physically teaching machines across different settings is time-consuming and labor-intensive. Simulation has long been seen as a solution, but prior physics engines have struggled to create sufficiently realistic and diverse virtual environments that capture the complexity of the real world. SceneSmith addresses this gap by leveraging AI agents powered by vision-language models—systems trained on vast amounts of text and images from the internet—to generate rich, detailed indoor spaces.
The three-agent design is elegant in how it distributes creative and quality-control tasks. The designer generates elements, the critic ensures practicality (for instance, flagging when a bathtub doesn't belong in a living room), and the orchestrator manages their collaboration. Over 200 user evaluations found the system's visuals more realistic than competing approaches 90 percent of the time, and it adhered more closely to text prompts. Critically, when researchers tested a robot policy trained entirely on real-world data in SceneSmith's generated environments—without the policy ever having seen such scenes—it successfully completed tasks like moving an apple from a bowl to a cutting board, suggesting the simulations are faithful enough to transfer real-world-trained skills.
For example, today's edition would include:
AI-summarized, only the topics you pick: one digest a day via Email, LINE, or Slack.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Momentic Inc. launched Mo, an AI agent that opens an app from a URL and prompt, spins up a swarm of agents to…
Anthropic launched Claude Sonnet 5.5, a mid-tier model for everyday tasks, priced at $2 per million input toke…
Oracle group VP Johnnie Konstantas told theCUBE that agentic AI lets agents and sub-agents act as a proxy for…
Qiagen has manually curated biomedical data for more than 25 years with over 150 MD- and PhD-level experts, Bh…
NEAR, the token of the NEAR Protocol, has more than doubled in value over the past two weeks, helped by surgin…

NVIDIA announced its Open Agent Safety Platform, made of the OpenShell open-source runtime and the Sentry refe…
