
What happened
The same requirements document was given to gpt-oss:120b (single-shot generation) and Claude Opus 5 (agent), both building a Next.js 14 + TypeScript + Prisma/SQLite weather app; independent inspection on 2026-10-01 added 3 new defects to each, bringing experiment #1 to 24 and experiment #2 to 7.
Why it matters
Both models produced defects that passed every automated gate — type checks, tests, builds — yet only surfaced when humans looked at the actual screen and data, suggesting machine gates alone are not enough for people who need to know what to trust a local model with.
What to watch
The results hinge on whether the local LLM can run as an agent, which the article says is a candidate for a future experiment. Watch for the second installment, which will cover what requirements could have prevented each defect layer.
WHO IT HITSDevelopers and engineering managers deciding which coding tasks can safely be handed to a local LLM for confidential work, and QA or review staff who need to know where human inspection remains irreplaceable.
Summaries like this, in your inbox every morning.
The experiment grew out of a practical concern: people handling confidential work would prefer to keep AI on a local machine. The question was how far a local LLM can be trusted. To find the boundary, the author had a cloud AI (Gemini) write a requirements document for a weather tracking app, then gave that same document to a local LLM and to a cloud AI for implementation. The two runs differed in form — single-shot generation versus an agent that could write files and run commands — and in prior knowledge, since the cloud AI had read the local LLM's experiment records. The author explicitly states these cannot be separated, so the comparison is not a ranking.
The defects fell into five layers. The first layer (format, config, types, schema) was largely caught by machine gates: the local LLM's output had 13 defect types found before gates, 12 caught by machines; the cloud AI had none in this layer. The remaining four layers — unstated assumptions, misreadings, unstated invariants, and UI/failure display — survived every gate on both sides. Both models displayed failures as non-failures: the local LLM silently substituted a meaningless count for an uncalculable metric, and the cloud AI showed 'not an error' alongside a fetch failure. Both failed to validate coordinate ranges, allowing latitude 999.
A mutation test added a concrete finding: when the time interpretation was switched to local time while keeping input validation, Tokyo and Los Angeles runs failed 10 tests each, but UTC runs failed none — meaning time zone bugs are invisible when the machine runs in UTC, common on servers and CI. The second installment will cover what requirements could have prevented and what they could not.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Carvana and CarMax are jockeying as agentic AI tools look to fix the worst parts of car buying, with AI shoppi…

Google DeepMind launched Gemini 4 Argon for coding, enterprise knowledge work and cyber defense, claiming firs…

Bloomberg reports Amazon's delivery smart glasses shoot still images at intervals during walks, possibly thous…

After forcing Hermes Agent's backend to Vulkan with the command "hermes config set local_runtime.backend vulka…

The Information reports Google's "AI Contribution Pilot Program" pays about 100 digital publishers, including…

Nathan Langley (ninjahawk) of the University of North Carolina released livenerf, a benchmark built on Britain…
