AIToday
Large Language ModelsAI Coding AssistantsAI Business & IndustryTomasz Tunguz (Theory Ventures)Published: Sep 18, 2026, 04:00 JST

UC Berkeley: harness cuts AI answer cost 71%

UC Berkeley: harness cuts AI answer cost 71%

3 Key Points

  1. What happened

    A UC Berkeley study found GPT-5.6 Sol costs 71% less on Pi than on Claude Code, returning the same result, and none of 42 harness comparisons showed a statistically significant quality difference.

  2. Why it matters

    If the same model returns the same result at a fraction of the cost, the system built around the model — not the model itself — becomes the source of cost advantage, according to the study.

  3. What to watch

    The finding covers 42 harness comparisons and one model; whether the pattern holds across other models and tasks is the test. The study says effective harnesses require deep customer understanding, relevant evals, and a factory for automating hill climbing.

WHO IT HITSThis lands hardest on AI product and engineering teams at software startups, who now face evidence that the layer they build around a model, not the model choice, drives gross margin. It also matters to investors pricing those startups on cost structure.

Not sure about something? Ask the AI

Summaries like this, in your inbox every morning.

Context & Analysis

The study's finding arrives as companies weigh how much to spend giving every task access to a state-of-the-art model. The article's example of two startups bidding on the same $250k contract — one calling a state-of-the-art model on every step, the other reserving it for the few steps that need it — shows how that choice flows straight to gross margin: 38% against 75%. The cheaper startup also pays back its sales cost in half the time, which the article frames as the difference between hiring twice as fast and not.

The article argues the cost advantage is hard to copy. Routing to a cheap model is only safe if you know which tasks it clears, and that knowledge comes from watching thousands of versions of the same work — the cost advantage and the moat are the same asset. The more expensive setup loses money on every evaluation past 8,600, forcing it to ration usage just as the customer finds it most valuable.

The article is careful not to promise 100% margins: as inference gets cheaper, buyers will ask more of their agents, shifting the equilibrium. What persists, it says, is the gap between the two companies. Whether that gap holds across models and tasks beyond the 42 comparisons is the open question.

FAQ
What is a harness in this context?
The article defines harnesses as the systems that control AI agents, coalescing common workflows into repeatable patterns of deterministic code or skills.
Did the cheaper harness produce worse answers?
No. None of the 42 harness comparisons showed a statistically significant quality difference, and the same model returned the same result.
What makes an effective harness?
The article lists a deep customer understanding, a collection of relevant evals, and a factory for automating hill climbing.
Tomasz Tunguz (Theory Ventures)Read Original Article

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Workiva Amplify: traceable AI agents needed for financeSiliconANGLE AI · 1h ago
  • Anthropic rebuilds Claude Code Projects for parallel agentsTHE DECODER · 1h ago
  • OpenAI reports six agent incidents, including 'no obligation to be subservient'Fortune AI · 1h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articlePika, valued at $470 million, launches faster AI video tool