AIToday
RoboticsOpen-Source AIExponential IndustryPublished: Sep 28, 2026, 01:00 JST

240 rollouts lift robot assembly from 24% to 88%

240 rollouts lift robot assembly from 24% to 88%

3 Key Points

  1. What happened

    Dream Machines fine-tuned Physical Intelligence's open-source π0.5 model on an industrial assembly task; 240 human recovery rollouts raised success from 24% to 88% with under two hours of training data, and Real Time Chunking settings reached ninety-eight percent insertion success.

  2. Why it matters

    Human corrections and data diversity beat raw demonstration volume for reliability, suggesting robotic automation teams can cut machine-training spend by targeting edge-case interventions rather than recording dozens of repetitive hours.

  3. What to watch

    The 88% figure comes from the team's physical evaluations, so whether it holds outside that setup is the test; watch how Real Time Chunking parameters and execution blending carry into other tasks.

WHO IT HITSRobotics engineers and automation teams building physical AI assembly systems benefit most, since the result shows that investing in edge-case human interventions can lift reliability far faster and more cheaply than collecting more repetitive demonstrations.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The result matters less as a single benchmark than as a signal about where robotics data budgets do the most work. The Dream Machines team collected teleoperation data across multiple desks using Gello controllers and virtual reality headsets, then found that data quality and visual diversity improved reliability faster than simply recording dozens of repetitive demonstration hours. That is the interesting part: the expensive thing in industrial automation has long been collecting more and more demonstration data, and this outcome points the other way.

The gains came from introducing human corrections into the training pipeline, with 240 recovery rollouts doing the heavy lifting. Full model fine-tuning outperformed low-rank adaptation methods on physical success rates, and Real Time Chunking settings from Physical Intelligence pushed final insertion success to ninety-eight percent during evaluations.

For companies developing robotic automation, the practical reading is that spending on edge-case interventions and execution blending parameters may buy more reliability than volume. Whether the 88% and ninety-eight percent figures carry into different tasks, hardware, or less controlled environments is the open question, since these came from the team's own physical evaluations.

FAQ
How much training data did it take to hit 88%?
Under two hours of total training data, with 240 recovery rollouts added to the pipeline.
What pushed insertion success to ninety-eight percent?
Real Time Chunking inference settings from Physical Intelligence during physical evaluations.
How did fine-tuning compare to other methods?
Full model fine-tuning produced much higher physical success rates than low-rank adaptation methods.
Exponential IndustryRead Original Article

Get the latest Robotics news every morning

For example, today's edition would include:

  • HomeBody: GPT-6 Astra runs a Unitree G1 kitchen tidy-upTHE DECODER · 4h ago
  • John Deere Rolls Out New AI AssistantTop Companies AI · 20h ago
  • McKinsey: Japan suppliers can cover almost all humanoid robot partsTop Companies AI · 20h ago

AI-summarized, only the topics you pick: one digest a day via Email, LINE, or Slack.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleAnthropic veterans eye remote US land if AI goes awry: WSJ