AIToday
Large Language ModelsRoboticsZenn AI/MLPublished: Oct 9, 2026, 22:00 JST

Decisions API gpt-6-luna drives SO-101 arm via repeated choices

Decisions API gpt-6-luna drives SO-101 arm via repeated choices

Using OpenAI's Decisions API, model gpt-6-luna, a developer built a pipeline that answers questions like "is the toy visible" or "is it left, center, or right" and uses them to move an SO-101 arm. Each decision took about 0.3–0.7 seconds.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The project builds directly on an earlier experiment in which the author drove an SO-101 arm with Claude and MCP without any robot-specific trained model. That worked — the arm grabbed an orange iguana toy — but each decision took several seconds and had to be repeated dozens of times, making the whole process slow. The new attempt replaces free-form model output with a typed API that returns probabilities for fixed questions, which the author says answers more stably and improves the ability to make the robot do nothing when confidence is low.

The write-up is candid about where the difficulty actually sits. Directly adding joint angles made the arm crawl along the table, because rotating the shoulder and elbow also changes the gripper's height. Switching to inverse kinematics fixed the concept but surfaced new problems: the wrist rolled from −75° to −145° as a side effect, the arm stalled 30cm from its base, a camera-mounted view made "image left" drift, and the arm once moved opposite to the intended direction because the mapping between image directions and table directions depended on how the camera was mounted. The author eventually removed the human from the loop by having the arm pose itself and measure its own camera orientation using the gripper as a marker, and by reading the real joint limits recorded by lerobot-calibrate rather than trusting the URDF.

The results so far are presented as a working pipeline rather than a completed task. The author notes that the run is being logged to runs/ so that successful attempts can later be used as imitation-learning data, and floats a future hybrid where a learned policy such as π0.5 handles fine positioning while the Decisions API is reserved for judgments like whether the gripper actually caught the target.

FAQ
What is the OpenAI Decisions API?
It is an API that takes images or text as input and returns typed answers to questions about them. It supports predicate (probability a condition is true), choice (one option plus each option's probability), and score. It returns no coordinates or prose.
How much does it cost and which model does it use?
At the time of writing it was in public beta, only the gpt-6-luna model was available, and pricing was $0.10 per 1 million input tokens, with no charge for output.
How fast is it compared with the previous approach?
The author reports about 0.3–0.7 seconds per decision, much faster than the earlier Claude × MCP method, where a single decision took several seconds.

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleOpenAI halts Dark Clark, Iranian fake-byline influence ops