
What happened
SAIL generates a trajectory with a VLM, runs it in a simulator, scores progress from video, and sends the corrected trajectory to the arm — 73% success versus 25% for one-shot generation.
Why it matters
Scoring where progress stalls lets the system fix only the failing section of a motion rather than guessing at the whole thing, and corrected successes can become training data.
What to watch
The real-arm test was open-loop, so it cannot correct a shifted object mid-motion, and it covered one task with six trials per method, so the plan hinges on whether the bench setup can be reconstructed in simulation.
WHO IT HITSRobotics engineers building VLM-driven pick-and-place systems would need a simulator that matches their real bench, plus a way to score progress mid-motion; the body's own open-loop limit means teams handling shifting objects still need a separate correction layer.
Summaries like this, in your inbox every morning.
The gap SAIL targets is between a plan that looks right and a motion that actually moves an object. A VLM can emit a sequence of gripper positions, orientations, and open/close states, and an inverse kinematics controller can reach each waypoint, yet still fail to grasp — because reaching a point is not the same as being able to grip there. SAIL closes that gap by running the candidate trajectory in a simulator, extracting frames from the resulting video, and having an evaluation VLM estimate how far each sub-task has progressed on a 0–100% scale. Those progress scores are tied back to waypoints, so the generation VLM receives a trajectory annotated with where progress stalled, not just a pass/fail label — letting it preserve high-scoring segments and revise low-scoring ones.
The search side is framed as Monte Carlo tree search where each node is an entire trajectory and each branch is a revision of it. In the same-15-candidate comparison reported by the authors, combining exploration with correction reached 65% average success across six tasks, against 51% for breadth-first generation without feedback and 37% for depth-first repeated correction. On the laptop-closing task at that budget, however, the breadth-first method scored higher — so the approach is not uniformly best.
For real deployment, the burden shifts to reconstruction: the body describes detecting objects with GroundingDINO, segmenting with SAM2, building a 3D point cloud from depth, and placing them in the simulator via camera calibration and ArUco markers. Because real execution is open-loop and was verified on one task with six trials per method, whether the loop pays off likely hinges on how faithfully a given bench can be mirrored in simulation and how often the task repeats.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Stratechery published a piece after OpenAI's Dev Day, saying the showcased product is, frankly, pretty confusi…

Electronic Library and NTT Docomo Business co-developed "ELNET AI" and will launch the full version on October…

KnowledgeSense said on September 30 that CodeSense, a Japanese-made AI agent, will ship within a few weeks, an…

Instinct, founded last year by 23-year-old Noah Shinn, raised $1 billion in a Series C, lifting its valuation…

Palo Alto Networks CEO Nikesh Arora let an unreleased Anthropic model, Mythos, probe his company's internal sy…

Liquid AI released d1, a decision model it says is the first to outperform Jev on Hugging Face's Decision Inde…
