
What happened
A developer tested whether ChatGPT would judge the same remote-work scenario differently when only the subject's gender and role changed. Across 200 sessions, the model refused to give a probability estimate for a female employee half the time, but not for a male employee or manager.
Why it matters
The results suggest the model's judgments may reflect systematic bias. Manager estimates averaged 11.54 percentage points lower than employee estimates, and the gender and role differences were statistically significant (adjusted p < 0.0001).
What to watch
The test was run on gpt-5.6-sol with high reasoning effort on 9/3/2026, and the author says the analysis was verified independently because 'we can't trust AI to even count correctly sometimes.' Whether other models show similar patterns remains an open question.
Ask the AI about this article →
Summaries like this, in your inbox every morning.
The experiment addresses a practical concern: whether AI language models, which are increasingly used in workplace decision-making, treat people differently based on gender or position. The author designed a simple scenario—a remote worker with frequent internet outages—and varied only the subject's gender and role. This isolates the model's bias from other confounding factors.
The findings reveal two distinct patterns. First, the model was more likely to refuse to give a numeric estimate when the subject was a female employee, suggesting a different default stance toward that group. Second, when it did answer, it judged managers more leniently than employees, with estimates averaging 11.54 percentage points lower. The gender-by-role interaction was significant, meaning the combined effect of gender and role was not just the sum of their individual effects.
The author emphasizes that the analysis was verified manually, acknowledging that 'we can't trust AI to even count correctly sometimes.' This underscores a broader challenge: as these models are deployed in real settings, their outputs—including subtle biases—can shape decisions. The results are specific to gpt-5.6-sol on a single date, so the generalizability across other models remains unclear.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Nvidia Corp. CEO Jensen Huang said artificial general intelligence has arrived, following OpenAI's launch of G…

Saudi Arabia's state-backed AI company HUMAIN, led by CEO Tareq Amin, is positioning itself as a neutral hub f…

Alibaba's research division released Qwen-Drive 1.0, an AI model that handles spatial perception, traffic Q&A…

Google DeepMind ran 100 autonomous LLM agents using Gemini 3.1 Pro on 71 math problems

Fukushima Prefecture ran a proof-of-concept in fiscal 2025 with 100 paid accounts for two generative AI servic…

Azoma, an Agentic Commerce Optimisation platform, published what brands and digital shelf teams should look fo…
