
What happened
A Stanford, Carnegie Mellon, UC Berkeley, and Microsoft Research team ran 6,800+ math, coding, and science tasks on 8 top AI models. In 32% of cases, the cheaper model's total cost was higher.
Why it matters
Cheaper models can 'overthink' and 'overact,' so a stronger model sometimes finishes faster and cheaper, meaning price alone may mislead buyers picking a model.
What to watch
The same model can vary up to 9.7x in cost across repeated runs, so the real test is whether pricing holds up across varied tasks, not a single benchmark.
WHO IT HITSTeams choosing which AI model to deploy for coding or data tasks may overpay when they pick on sticker price alone. Procurement and product managers comparing model costs should weigh per-task outcomes, not headline rates.
Summaries like this, in your inbox every morning.
The study, run by researchers from Stanford, Carnegie Mellon, UC Berkeley, and Microsoft Research, tested 8 cutting-edge AI models on over 6,800 tasks spanning math, programming, and science. Its central finding is counterintuitive: a lower-priced model does not automatically translate to a lower bill. In one example, the pricier Gemini 3.1 Pro completed a YouTube frame-by-frame analysis task in 85 steps for 1 dollar, while the cheaper Gemini 3 Flash ran over 1,000 steps, failed to finish, and cost 14 dollars.
The researchers attribute this to AI 'overthinking' and 'overacting.' On complex tasks, weaker models tend to burn more reasoning steps and actions, so a stronger model can finish faster and cheaper. A separate example saw Gemini 3 Flash consume over 60,000 thinking tokens while the more capable GPT-5.4 solved the same problem in 25 tokens.
The findings also show that cost is not even stable within a single model: running the same instruction repeatedly produced a most-expensive run 9.7 times the cheapest. Lead author Lingzhao Chen's point that price alone should not decide which model is cheap suggests the real stakes fall on teams comparing AI costs, whose budgeting may hinge on whether they evaluate per-task outcomes rather than sticker rates.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Ghost AI raised $11 million, led by Andreessen Horowitz with Abstract, Audacious Ventures, Nova and SV Angel…
Moonshot AI closed its final private round at about a $50 billion valuation and is aiming for a Hong Kong IPO…

OpenAI told Fortune it is piloting a 'mission interview' for candidates on certain teams, starting with market…

Reactor said it added Nvidia's venture capital arm and Sapphire Ventures to its cap table, bringing its total…

OpenAI will roll out invisible text watermarks to all ChatGPT and Codex plan users in the EU within weeks, and…

Reflection announced Beam, a 501B-parameter open-weight model with 23B active parameters, claiming parity with…
