AIToday
Large Language ModelsRoboticsQiita 機械学習Published: Sep 30, 2026, 19:00 JST

Sakana AI's SAIL lifts robot task success to 73% via simulation rehearsal

Sakana AI's SAIL lifts robot task success to 73% via simulation rehearsal

3 Key Points

  1. What happened

    SAIL generates a trajectory with a VLM, runs it in a simulator, scores progress from video, and sends the corrected trajectory to the arm — 73% success versus 25% for one-shot generation.

  2. Why it matters

    Scoring where progress stalls lets the system fix only the failing section of a motion rather than guessing at the whole thing, and corrected successes can become training data.

  3. What to watch

    The real-arm test was open-loop, so it cannot correct a shifted object mid-motion, and it covered one task with six trials per method, so the plan hinges on whether the bench setup can be reconstructed in simulation.

WHO IT HITSRobotics engineers building VLM-driven pick-and-place systems would need a simulator that matches their real bench, plus a way to score progress mid-motion; the body's own open-loop limit means teams handling shifting objects still need a separate correction layer.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The gap SAIL targets is between a plan that looks right and a motion that actually moves an object. A VLM can emit a sequence of gripper positions, orientations, and open/close states, and an inverse kinematics controller can reach each waypoint, yet still fail to grasp — because reaching a point is not the same as being able to grip there. SAIL closes that gap by running the candidate trajectory in a simulator, extracting frames from the resulting video, and having an evaluation VLM estimate how far each sub-task has progressed on a 0–100% scale. Those progress scores are tied back to waypoints, so the generation VLM receives a trajectory annotated with where progress stalled, not just a pass/fail label — letting it preserve high-scoring segments and revise low-scoring ones.

The search side is framed as Monte Carlo tree search where each node is an entire trajectory and each branch is a revision of it. In the same-15-candidate comparison reported by the authors, combining exploration with correction reached 65% average success across six tasks, against 51% for breadth-first generation without feedback and 37% for depth-first repeated correction. On the laptop-closing task at that budget, however, the breadth-first method scored higher — so the approach is not uniformly best.

For real deployment, the burden shifts to reconstruction: the body describes detecting objects with GroundingDINO, segmenting with SAM2, building a 3D point cloud from depth, and placing them in the simulator via camera calibration and ArUco markers. Because real execution is open-loop and was verified on one task with six trials per method, whether the loop pays off likely hinges on how faithfully a given bench can be mirrored in simulation and how often the task repeats.

FAQ
Which VLM does SAIL use for generating and evaluating trajectories?
Both the generation side and the evaluation side use Gemini Robotics-ER 1.5, according to the experiments. SAIL is the surrounding mechanism that uses it to improve trajectories.
How well did SAIL work when moving from simulator to a real robot arm?
With a LeRobot SO-101 arm placing a blue block into a red bowl, SAIL searched up to 15 candidates in a reconstructed environment and succeeded in 5 of 6 trials. The authors attributed the remaining failure to position and orientation estimation error and insufficient reproduction of real contact behavior.
What did SAIL do with the successful trajectories it found?
It collected 120 successful trajectories in simulation and used them to train an Action Chunking with Transformers (ACT) model by imitation learning. On the real arm, this learned policy also succeeded in 5 of 6 trials, with the paper reporting an average execution time of 72.306 seconds versus 644.72 seconds for the MCTS method.
Qiita 機械学習Read Original Article

AI news that matters for your work, in one minute a day

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleLiquid AI's d1 tops Jev on Decision Index