AIToday
Large Language ModelsAI Business & IndustryGIGAZINE AIPublished: Oct 6, 2026, 13:00 JST

Cheap AI models cost more in 32% of tasks, study finds

Cheap AI models cost more in 32% of tasks, study finds

3 Key Points

  1. What happened

    A Stanford, Carnegie Mellon, UC Berkeley, and Microsoft Research team ran 6,800+ math, coding, and science tasks on 8 top AI models. In 32% of cases, the cheaper model's total cost was higher.

  2. Why it matters

    Cheaper models can 'overthink' and 'overact,' so a stronger model sometimes finishes faster and cheaper, meaning price alone may mislead buyers picking a model.

  3. What to watch

    The same model can vary up to 9.7x in cost across repeated runs, so the real test is whether pricing holds up across varied tasks, not a single benchmark.

WHO IT HITSTeams choosing which AI model to deploy for coding or data tasks may overpay when they pick on sticker price alone. Procurement and product managers comparing model costs should weigh per-task outcomes, not headline rates.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The study, run by researchers from Stanford, Carnegie Mellon, UC Berkeley, and Microsoft Research, tested 8 cutting-edge AI models on over 6,800 tasks spanning math, programming, and science. Its central finding is counterintuitive: a lower-priced model does not automatically translate to a lower bill. In one example, the pricier Gemini 3.1 Pro completed a YouTube frame-by-frame analysis task in 85 steps for 1 dollar, while the cheaper Gemini 3 Flash ran over 1,000 steps, failed to finish, and cost 14 dollars.

The researchers attribute this to AI 'overthinking' and 'overacting.' On complex tasks, weaker models tend to burn more reasoning steps and actions, so a stronger model can finish faster and cheaper. A separate example saw Gemini 3 Flash consume over 60,000 thinking tokens while the more capable GPT-5.4 solved the same problem in 25 tokens.

The findings also show that cost is not even stable within a single model: running the same instruction repeatedly produced a most-expensive run 9.7 times the cheapest. Lead author Lingzhao Chen's point that price alone should not decide which model is cheap suggests the real stakes fall on teams comparing AI costs, whose budgeting may hinge on whether they evaluate per-task outcomes rather than sticker rates.

FAQ
Is the cheaper AI model always the more affordable choice?
No. In 32% of cases studied, the low-cost model ended up with a higher total cost. This happened because weaker models took more steps to solve the same task.
Can the same AI model give different costs for the same task?
Yes. When the team ran the same instruction on the same model repeatedly, the most expensive run was 9.7 times the cheapest run.
What did the lead author say about judging model prices?
Lead author Lingzhao Chen said you should not decide which model is cheap by looking at price alone.

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleNetApp's Novus hits 100TB/s to ease AI metadata bottlenecks