
What happened
A UC Berkeley study found GPT-5.6 Sol costs 71% less on Pi than on Claude Code, returning the same result, and none of 42 harness comparisons showed a statistically significant quality difference.
Why it matters
If the same model returns the same result at a fraction of the cost, the system built around the model — not the model itself — becomes the source of cost advantage, according to the study.
What to watch
The finding covers 42 harness comparisons and one model; whether the pattern holds across other models and tasks is the test. The study says effective harnesses require deep customer understanding, relevant evals, and a factory for automating hill climbing.
WHO IT HITSThis lands hardest on AI product and engineering teams at software startups, who now face evidence that the layer they build around a model, not the model choice, drives gross margin. It also matters to investors pricing those startups on cost structure.
Summaries like this, in your inbox every morning.
The study's finding arrives as companies weigh how much to spend giving every task access to a state-of-the-art model. The article's example of two startups bidding on the same $250k contract — one calling a state-of-the-art model on every step, the other reserving it for the few steps that need it — shows how that choice flows straight to gross margin: 38% against 75%. The cheaper startup also pays back its sales cost in half the time, which the article frames as the difference between hiring twice as fast and not.
The article argues the cost advantage is hard to copy. Routing to a cheap model is only safe if you know which tasks it clears, and that knowledge comes from watching thousands of versions of the same work — the cost advantage and the moat are the same asset. The more expensive setup loses money on every evaluation past 8,600, forcing it to ration usage just as the customer finds it most valuable.
The article is careful not to promise 100% margins: as inference gets cheaper, buyers will ask more of their agents, shifting the equilibrium. What persists, it says, is the gap between the two companies. Whether that gap holds across models and tasks beyond the 42 comparisons is the open question.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Christie Kozlik, chief accounting officer at Accel Entertainment Inc., told theCUBE that finance adoption is g…
At Workiva's Amplify event, theCUBE Research analyst Krista Case said firms must trace where AI insights came…
Pika, valued at $470 million, launched a program that automates which AI model creates each video element and…

Anthropic rebuilt the Projects feature in Claude Code

Zuckerberg pushed back on coordinated AI slowdown calls on September 15, saying Meta delayed its Muse agent fo…

OpenAI published a voluntary disclosure framework and reported six misaligned agent incidents, including an As…
