AIToday
Large Language ModelsHacker NewsPublished: Apr 26, 2026, 22:00 JST1 min read

Research reveals GPT, Gemini, and Claude all hallucinate when counting large datasets—but a structured protocol called KIS fixes the problem

3 Key Points

  1. A peer-reviewed study (published April 26, 2026) tested three AI models—GPT-5.3 Instant, Gemini 3 Flash, and Claude Sonnet 4.6—on counting tasks with 200 to 2,000 items. Results: Gemini overcounted by 38 items at 1,000 entries under baseline conditions; GPT abandoned the task entirely beyond 800 items; Claude stayed accurate without any help. All three models made counting errors that looked like confident answers.

  2. Each model fails in a different way. Gemini makes up plausible-sounding numbers (called 'confabulation'); GPT refuses to try on large datasets (called 'avoidance'); Claude hides its reasoning process so errors go undetected (called 'process-opaque'). Using Chain-of-Thought prompting (a technique to make AI explain its steps) actually made GPT worse, triggering false counts even on small datasets of 200 items.

  3. When researchers applied KIS—a structured protocol that separates counting, verification, and reporting into distinct logged steps—all three models achieved 100% accuracy across all dataset sizes. For anyone deploying AI in accounting, legal discovery, supply chain audits, or any field where miscounts have legal or financial consequences, this means: use structured protocols and demand audit trails, or your AI's confident answer could be quietly wrong.

Ask the AI about this article →

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Meta's Muse Spark 1.3 matches GPT-5.6-Sol, claims #3 modelLatent Space · 4h ago
  • CABiNet vs YOLO26-sem on UAVid: Accuracy, Compute, and GPU Latencyr/MachineLearning · 4h ago
  • C++ PCN library nears backprop accuracyr/MachineLearning · 4h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleBerenberg initiates Buy on Palo Alto Networks at $215 target; Evercore upgrades Arista Networks ahead of May 5 earnings; AMD upgraded — analysts see AI as cybersecurity tailwind, not threat