
Mathematician Adam Kucharski tested Microsoft Copilot in "Auto" mode with identical datasets labeled with different countries (UK, US, France, Germany, Italy). The tool generated country-specific differences that did not exist in the data—for example, claiming Italians were three times more likely to show interest in arts careers than Brits, and Americans were 1.5 times more business-oriented than the French.
Fast models (Copilot Auto, Gemini Flash 3.5) relied on stereotypes baked into their language models rather than actually analyzing the data. ChatGPT Instant and Claude Opus 4.7 automatically switched to extended reasoning mode, wrote Python code to analyze the dataset, and correctly identified that the data was identical.
The default "Auto" mode in Copilot is the main issue, since most users stick with the default setting. Real-world risk: analyses applied to actual datasets where groups have no real differences could appear worlds apart due to the model's built-in assumptions about demographic groups.
Ask the AI about this article →
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
CBTS Technology Solutions LLC launched Forge Agents, a platform that turns a plain-language job description in…
Imec CEO Patrick Vandenameele said at SEMICON Taiwan 2026 that the Belgian research center is broadening its c…

Alphabet's AI Overviews now reach over 2.5 billion monthly users through Google Search, and its ad business ge…

Amazon Web Services (AWS) has integrated its fully managed data warehouse service, Amazon Redshift, with Agent…

Visual Studio Code 1.135 now includes an experimental 'Rubber Duck' feature that lets developers request a sec…

Anthropic is making a permanent 25% increase to the usage limits in Claude Code, effective after September 14
