A peer-reviewed study (published April 26, 2026) tested three AI models—GPT-5.3 Instant, Gemini 3 Flash, and Claude Sonnet 4.6—on counting tasks with 200 to 2,000 items. Results: Gemini overcounted by 38 items at 1,000 entries under baseline conditions; GPT abandoned the task entirely beyond 800 items; Claude stayed accurate without any help. All three models made counting errors that looked like confident answers.
Each model fails in a different way. Gemini makes up plausible-sounding numbers (called 'confabulation'); GPT refuses to try on large datasets (called 'avoidance'); Claude hides its reasoning process so errors go undetected (called 'process-opaque'). Using Chain-of-Thought prompting (a technique to make AI explain its steps) actually made GPT worse, triggering false counts even on small datasets of 200 items.
When researchers applied KIS—a structured protocol that separates counting, verification, and reporting into distinct logged steps—all three models achieved 100% accuracy across all dataset sizes. For anyone deploying AI in accounting, legal discovery, supply chain audits, or any field where miscounts have legal or financial consequences, this means: use structured protocols and demand audit trails, or your AI's confident answer could be quietly wrong.
Ask the AI about this article →
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Meta released Muse Spark 1.3, which per AAII is now the #3 model in the world

The original author of CABiNet (ICRA 2021) rebuilt the repo and compared its performance against YOLO26-sem on…

A developer built Deepity, a C++ machine learning library, to test Predictive Coding Networks (PCNs), an alter…

Meta Platforms Inc. released Muse Spark 1.3, which it says puts it at the same level as the most recent models…
Meta announced Muse Spark 1.3 on September 2 and started offering it to developers through Muse Code and the M…

The author tested AI writing tools over three years and found AI automates mechanical tasks but does not reduc…
