AIToday
Large Language ModelsRoboticsRobohubPublished: Aug 7, 2026, 19:00 JST

MIT system uses AI agents to create realistic virtual training spaces for robots

MIT system uses AI agents to create realistic virtual training spaces for robots

3 Key Points

  1. What happened

    Researchers at MIT CSAIL and Toyota Research Institute developed SceneSmith, a system that uses three AI agents powered by GPT-5.2 to generate detailed 3D indoor environments. The agents—a designer, critic, and orchestrator—work together to create realistic scenes with up to six times more objects per space than prior methods, which robots can then practice tasks in before real-world deployment.

  2. Why it matters

    Creating training data for robots is labor-intensive and time-consuming. By generating rich, physics-accurate virtual environments from text prompts alone, SceneSmith lets engineers test robot behavior in simulation before deployment, saving real-world trial-and-error. In evaluation, a robot policy trained on real-world data successfully completed tasks in SceneSmith scenes—suggesting the simulations are realistic enough to transfer learning.

  3. What to watch

    The system currently takes multiple hours to generate a single scene due to the agents' detailed scrutiny of each object. Researchers note that increased computing power could dramatically improve efficiency, and they plan to expand to deformable objects like sponges if 3D libraries become available. The team presented findings as a spotlight at the 2026 International Conference on Machine Learning.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The core challenge roboticists face is that robots learn best through experience, but physically teaching machines across different settings is time-consuming and labor-intensive. Simulation has long been seen as a solution, but prior physics engines have struggled to create sufficiently realistic and diverse virtual environments that capture the complexity of the real world. SceneSmith addresses this gap by leveraging AI agents powered by vision-language models—systems trained on vast amounts of text and images from the internet—to generate rich, detailed indoor spaces.

The three-agent design is elegant in how it distributes creative and quality-control tasks. The designer generates elements, the critic ensures practicality (for instance, flagging when a bathtub doesn't belong in a living room), and the orchestrator manages their collaboration. Over 200 user evaluations found the system's visuals more realistic than competing approaches 90 percent of the time, and it adhered more closely to text prompts. Critically, when researchers tested a robot policy trained entirely on real-world data in SceneSmith's generated environments—without the policy ever having seen such scenes—it successfully completed tasks like moving an apple from a bowl to a cutting board, suggesting the simulations are faithful enough to transfer real-world-trained skills.

FAQ
How does SceneSmith create 3D scenes?
Three AI agents working with GPT-5.2 create scenes in stages: a designer agent generates the layout and objects, a critic agent reviews for realism and practicality, and an orchestrator agent manages their back-and-forth. Once approved, the scene is loaded directly into physics simulation software.
How many objects can SceneSmith include in a scene?
SceneSmith creates environments with up to six times more items per scene than prior methods such as HSM and Holodeck. The system can also generate articulated objects like cabinets that robots can open and close, which prior systems rarely included.
How long does it take to generate a scene?
It can take multiple hours to produce a single scene because the agents are creating and closely scrutinizing each object. Researchers note that more computing power could dramatically increase efficiency.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Agentic AI pushes identity to front of data security, Oracle saysSiliconANGLE AI · 1h ago
  • Momentic launches Mo, an AI agent that tests apps without scriptsSiliconANGLE AI · 1h ago
  • Anthropic debuts Claude Sonnet 5.5, 30% fasterSiliconANGLE AI · 1h ago

AI-summarized, only the topics you pick: one digest a day via Email, LINE, or Slack.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleCisco Targets $9B AI Orders, Up from $5B—Stock Up 60% in 2026