AIToday
Large Language ModelsZenn AI/MLPublished: Oct 1, 2026, 10:00 JST

Local LLM vs Claude Opus 5: same app, same blind spots

Local LLM vs Claude Opus 5: same app, same blind spots

3 Key Points

  1. What happened

    The same requirements document was given to gpt-oss:120b (single-shot generation) and Claude Opus 5 (agent), both building a Next.js 14 + TypeScript + Prisma/SQLite weather app; independent inspection on 2026-10-01 added 3 new defects to each, bringing experiment #1 to 24 and experiment #2 to 7.

  2. Why it matters

    Both models produced defects that passed every automated gate — type checks, tests, builds — yet only surfaced when humans looked at the actual screen and data, suggesting machine gates alone are not enough for people who need to know what to trust a local model with.

  3. What to watch

    The results hinge on whether the local LLM can run as an agent, which the article says is a candidate for a future experiment. Watch for the second installment, which will cover what requirements could have prevented each defect layer.

WHO IT HITSDevelopers and engineering managers deciding which coding tasks can safely be handed to a local LLM for confidential work, and QA or review staff who need to know where human inspection remains irreplaceable.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The experiment grew out of a practical concern: people handling confidential work would prefer to keep AI on a local machine. The question was how far a local LLM can be trusted. To find the boundary, the author had a cloud AI (Gemini) write a requirements document for a weather tracking app, then gave that same document to a local LLM and to a cloud AI for implementation. The two runs differed in form — single-shot generation versus an agent that could write files and run commands — and in prior knowledge, since the cloud AI had read the local LLM's experiment records. The author explicitly states these cannot be separated, so the comparison is not a ranking.

The defects fell into five layers. The first layer (format, config, types, schema) was largely caught by machine gates: the local LLM's output had 13 defect types found before gates, 12 caught by machines; the cloud AI had none in this layer. The remaining four layers — unstated assumptions, misreadings, unstated invariants, and UI/failure display — survived every gate on both sides. Both models displayed failures as non-failures: the local LLM silently substituted a meaningless count for an uncalculable metric, and the cloud AI showed 'not an error' alongside a fetch failure. Both failed to validate coordinate ranges, allowing latitude 999.

A mutation test added a concrete finding: when the time interpretation was switched to local time while keeping input validation, Tokyo and Los Angeles runs failed 10 tests each, but UTC runs failed none — meaning time zone bugs are invisible when the machine runs in UTC, common on servers and CI. The second installment will cover what requirements could have prevented and what they could not.

FAQ
What models were tested?
gpt-oss:120b running on Ollama on a Mac Studio, and Claude Opus 5. The local model used single-shot generation, while the cloud model ran as an agent that could write files and execute commands.
Did both models have the same starting conditions?
No. The cloud AI read the local LLM's experiment records before implementing, so it already knew about pitfalls found earlier. The article also notes the defect-counting methods differed and that model differences cannot be separated from form differences.
What defects did the local LLM have that machines missed?
It displayed the total snapshot count (72 minus 1 = 71) as the 'forecast change count' even when the forecast never changed, showed a 9-hour time zone offset (screen said 17:30 while actual time was 02:31), and misread the 8:00 baseline requirement so the scoring function could not work.

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleGoogle unveils Gemini 4 Argon, limits early access