
What happened
Earn an Honest Dollar tested 16 AI models on 42 page pairs with decoy data. Without the "don't guess" instruction, 405 of 573 missing fields (70.7%) got fabricated values; with it, 116 of 574 (20.2%).
Why it matters
Every one of the 16 models fabricated fewer values when told to return null instead of guessing, suggesting the behavior can be curbed without retraining the models themselves.
What to watch
The results are limited to web data extraction, and Earn an Honest Dollar says other uses need further testing. Its check experiment with GPT-6 Luna caught 38 of 49 fabricated values without wrongly rejecting the 47 correct ones.
WHO IT HITSBusinesses and teams that use AI agents to pull data from web pages — such as product listings, prices, or author names — may see fewer invented values simply by adding a no-guessing instruction, though the finding is so far limited to web extraction. Because one AI agent may not verify what another returns, the risk of silently fabricated data is a live concern for anyone automating data collection.
Summaries like this, in your inbox every morning.
The experiment grew out of a specific worry for Earn an Honest Dollar, which runs a marketplace where AI agents buy and sell services from one another. The buying agent cannot always verify what it receives, so an agent that quietly fills in missing values — say, a price that was never listed — is a serious problem rather than a cosmetic one.
To test this, the company built 42 test pages across seven page formats, each pair nearly identical except that one version contained the sought-after information and the other had it removed. Deliberate traps were placed inside: a product page with a stale "493ドル" price but no current one, or an article carrying the name "Omar Tam" as its fact-checker where no author existed. Across 16 models, the no-guessing instruction cut fabrication sharply, and the old-price trap proved the clearest illustration — all 16 models reported 493ドル as the current price without the instruction, but only one did with it.
The company also tried a second layer of defense, having a cheaper AI verify whether extracted values had support in the original page. It notes that its benchmark covers only extraction from the web, and that whether the same pattern holds for other uses will require more testing. For teams automating data collection, the practical stake may be whether a single added instruction, rather than a rebuilt model, is enough to make AI output trustworthy — a question whose answer likely depends on the task at hand. Whether a checking model proves reliable at scale is another open point.
For example, today's edition would include:
AI-summarized, only the topics you pick: one digest a day via Email, LINE, or Slack.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Nvidia's DGX Spark, launched at CES 2025 as Project DIGITS, began shipping in October, with VP Adel El Hallak…

Google said its Gemini 3.8 Live voice model now has a Live Avatar feature, and that Live Avatar-equipped Gemin…

U.K. AI minister Kanishka Narayan said nations must "harden and build your defenses" against AI risk, after ex…

On the 2,000-question typed-decisions benchmark, Supersonic Labs' free Julia 1 scored 73.15% vs

Fireworks AI announced Ember-1, a Kimi K3-based model built for its project to create specialized models devel…

Dymocks Education is closing its classrooms and urging parents to use ChatGPT or Gemini instead, with CEO Mark…
