
MIT researchers have created SceneSmith, a system that uses three AI agents to generate highly realistic 3D indoor environments for robot training.
Powered by the vision-language model GPT-5.2, the system creates scenes with far more objects and detail than prior methods, allowing robots to practice tasks in simulation before deployment in the real world.
In tests, robots trained on real-world data successfully completed tasks in these simulated spaces, indicating the environments are realistic enough to be useful for training.
What happened
Researchers at MIT CSAIL and Toyota Research Institute developed SceneSmith, a system that uses three AI agents powered by GPT-5.2 to generate detailed 3D indoor environments. The agents—a designer, critic, and orchestrator—work together to create realistic scenes with up to six times more objects per space than prior methods, which robots can then practice tasks in before real-world deployment.
Why it matters
Creating training data for robots is labor-intensive and time-consuming. By generating rich, physics-accurate virtual environments from text prompts alone, SceneSmith lets engineers test robot behavior in simulation before deployment, saving real-world trial-and-error. In evaluation, a robot policy trained on real-world data successfully completed tasks in SceneSmith scenes—suggesting the simulations are realistic enough to transfer learning.
What to watch
The system currently takes multiple hours to generate a single scene due to the agents' detailed scrutiny of each object. Researchers note that increased computing power could dramatically improve efficiency, and they plan to expand to deformable objects like sponges if 3D libraries become available. The team presented findings as a spotlight at the 2026 International Conference on Machine Learning.
Robots are increasingly visible in public spaces, but they remain limited as practical assistants for kitchens, factories, and other real-world settings. A major bottleneck is training data: teaching robots to handle diverse tasks across varied environments requires either extensive physical interaction or vast amounts of manually captured real-world data. Russ Tedrake, the Toyota Professor of Electrical Engineering and Computer Science at MIT and a principal investigator at MIT CSAIL, identifies the core problem: "While there has been significant progress over the last few years in the physics engines that power robotics simulators, one of the remaining challenges has been creating sufficiently rich and diverse simulation content to capture the complexity of the real world."
SceneSmith, developed by researchers at MIT CSAIL and Toyota Research Institute, tackles this challenge using an agentic approach. The system leverages three AI agents, each powered by GPT-5.2, a state-of-the-art vision-language model trained on internet-scale text and images. The agents operate in distinct roles: a designer generates scene elements, a critic evaluates realism and practicality, and an orchestrator manages their back-and-forth dialogue. According to MIT EECS PhD student Nicholas Pfaff, a lead researcher on the project, "We've found that the system can construct 3D scenes the way a human designer would." The team created over 1,300 scenes, which Pfaff notes made "insanely creative and diverse arrangements" that emerged without explicit instruction: "I hadn't taught the system to do that in the prompts; it just improvised."
The resulting virtual environments are orders of magnitude richer than prior methods. Users can prompt SceneSmith to generate specific scenes—"generate a garage with a car, a workbench, tires stacked in the corner, and a ladder against the wall"—and receive detailed, interactive 3D spaces with up to six times more objects per scene than baseline approaches. Unlike prior systems, SceneSmith includes articulated objects like cabinets that robots can open and close, expanding the range of skills robots can practice. When researchers tested a robot policy trained on real-world data in these simulated spaces, it successfully completed manipulation tasks such as placing an apple on a cutting board. Humans agreed with the model's evaluations of robot failures over 99 percent of the time, validating the system as a filter for flawed approaches before deployment. Over 200 user evaluations found the system's visuals more realistic than competing methods 90 percent of the time.
The practical payoff is significant: engineers can evaluate and refine robot behavior in simulation rather than through extensive physical trials. However, SceneSmith comes with a computational cost. Generating a single scene can take multiple hours because the agents scrutinize each object in detail. Researchers believe more computing power could dramatically improve speed, and they plan to expand the system to handle deformable objects like sponges if suitable 3D libraries become available. Jeremy Binagia, an applied scientist at Amazon Robotics, calls SceneSmith "a significant advance," noting it "advances the state of the art in several ways, including pushing the limits of the density of objects in the simulated environment" and "creating assets that are not constrained to a fixed library." The team presented their findings as a spotlight presentation at the 2026 International Conference on Machine Learning, supported in part by Amazon, the U.S. Office of Naval Research, the Toyota Research Institute, and the U.S. National Science Foundation.
The core challenge roboticists face is that robots learn best through experience, but physically teaching machines across different settings is time-consuming and labor-intensive. Simulation has long been seen as a solution, but prior physics engines have struggled to create sufficiently realistic and diverse virtual environments that capture the complexity of the real world. SceneSmith addresses this gap by leveraging AI agents powered by vision-language models—systems trained on vast amounts of text and images from the internet—to generate rich, detailed indoor spaces.
The three-agent design is elegant in how it distributes creative and quality-control tasks. The designer generates elements, the critic ensures practicality (for instance, flagging when a bathtub doesn't belong in a living room), and the orchestrator manages their collaboration. Over 200 user evaluations found the system's visuals more realistic than competing approaches 90 percent of the time, and it adhered more closely to text prompts. Critically, when researchers tested a robot policy trained entirely on real-world data in SceneSmith's generated environments—without the policy ever having seen such scenes—it successfully completed tasks like moving an apple from a bowl to a cutting board, suggesting the simulations are faithful enough to transfer real-world-trained skills.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
Amazon and Google are intensifying competitive efforts against The Trade Desk (TTD), a major digital advertisi…

OpenAI introduced Premium Seats for ChatGPT Business, priced at $125 per user per month ($100 with annual bill…

Computer scientists at University of Tübingen, Max Planck Institute, MATS Research, and Snyk discovered a meth…

Anthropic pledged to embed machine-readable watermarks in Claude-generated text and digitally signed provenanc…

Anthropic has signed the EU AI Act Code of Practice and will embed invisible watermarks in Claude-generated te…

Anthropic has agreed to pay $9.1 billion over 20 years to Riot Platforms Inc., a Bitcoin miner turned data cen…

The AI news that matters, in one minute each morning.
Sign up free