
What happened
DrivingBench, run by three engineers at the AI startup Axiom Math, connected GPT-6 Astra, Claude Fable 5.1, Grok 4.6 and GPT-5.6 Sol to a 2022 Toyota Corolla. Only GPT-6 Astra completed the roughly 135 m course, in 5 minutes 22 seconds.
Why it matters
Drivers were told to brake at any time and the vehicle was speed-limited through a cone-lined course, so the result is a benchmark of physical capability, not a usable self-driving system.
What to watch
The team is considering DrivingBench v2, with more trials per model and harder courses, so whether the outcome holds depends on whether the renamed sandbox still works at greater scale.
WHO IT HITSThe result lands on robotics and autonomous-systems teams weighing whether general-purpose AI models can be trusted with physical hardware, and it flags response latency as an evaluation criterion alongside decision accuracy: while the AI is still thinking, the car keeps moving from the last command.
Summaries like this, in your inbox every morning.
DrivingBench was not set up to build a practical self-driving system. Its three founders — Aditya Ramabadran, Tobias Gessler and Simon Marns, who work at the AI startup Axiom Math — built it after seeing recent demos of AI drawing pictures or moving objects with robots, and wondering whether general-purpose AI could already drive. The test car was a 2022 Toyota Corolla fitted with the in-car device 'comma four' and the open-source driver-assistance software 'openpilot', with a human driver seated and ready to brake.
The instructions given to the AI ran under 600 words: stay inside a reverse U-shaped parking-lot course marked by cones, then park in a space marked by a small blue cone. Increasing distance was the main goal, with a fast finish requested where possible. Two problems surfaced ahead of driving skill itself. Models that recognized they were operating a real car refused on safety grounds, with GPT-6 Astra refusing especially often, forcing hours of adjusting prompts and system names; and because the car keeps moving on the previous command while the AI considers its next move, slower responses raise the risk of leaving the course. DrivingBench says curve failures came down mainly to 'perception' — reading the correct driving position from the arrangement of cones.
Marns described the result as an interesting experiment showing a new capability emerging in AI, and argued such tests have value because AI long used mainly inside computer screens is likely to reach the physical world more often through robots. Whether that reading holds may depend on whether the team's planned DrivingBench v2, with more trials per model, different inference settings and harder courses, reproduces the result.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Stratechery published a piece after OpenAI's Dev Day, saying the showcased product is, frankly, pretty confusi…

President Trump and the CEOs of the most powerful AI companies—including Elon Musk, Mark Zuckerberg, Dario Amo…

President Trump and executives including Mark Zuckerberg, Greg Brockman, Jensen Huang, and Elon Musk signed an…

Electronic Library and NTT Docomo Business co-developed "ELNET AI" and will launch the full version on October…

KnowledgeSense said on September 30 that CodeSense, a Japanese-made AI agent, will ship within a few weeks, an…

Instinct, founded last year by 23-year-old Noah Shinn, raised $1 billion in a Series C, lifting its valuation…
