AIToday
Large Language ModelsAI Safety & AlignmentHacker NewsPublished: Sep 7, 2026, 22:00 JST2 min read

ChatGPT judges female employees more harshly, study finds

ChatGPT judges female employees more harshly, study finds

3 Key Points

  1. What happened

    A developer tested whether ChatGPT would judge the same remote-work scenario differently when only the subject's gender and role changed. Across 200 sessions, the model refused to give a probability estimate for a female employee half the time, but not for a male employee or manager.

  2. Why it matters

    The results suggest the model's judgments may reflect systematic bias. Manager estimates averaged 11.54 percentage points lower than employee estimates, and the gender and role differences were statistically significant (adjusted p < 0.0001).

  3. What to watch

    The test was run on gpt-5.6-sol with high reasoning effort on 9/3/2026, and the author says the analysis was verified independently because 'we can't trust AI to even count correctly sometimes.' Whether other models show similar patterns remains an open question.

Ask the AI about this article →

Summaries like this, in your inbox every morning.

Context & Analysis

The experiment addresses a practical concern: whether AI language models, which are increasingly used in workplace decision-making, treat people differently based on gender or position. The author designed a simple scenario—a remote worker with frequent internet outages—and varied only the subject's gender and role. This isolates the model's bias from other confounding factors.

The findings reveal two distinct patterns. First, the model was more likely to refuse to give a numeric estimate when the subject was a female employee, suggesting a different default stance toward that group. Second, when it did answer, it judged managers more leniently than employees, with estimates averaging 11.54 percentage points lower. The gender-by-role interaction was significant, meaning the combined effect of gender and role was not just the sum of their individual effects.

The author emphasizes that the analysis was verified manually, acknowledging that 'we can't trust AI to even count correctly sometimes.' This underscores a broader challenge: as these models are deployed in real settings, their outputs—including subtle biases—can shape decisions. The results are specific to gpt-5.6-sol on a single date, so the generalizability across other models remains unclear.

FAQ
How was the test conducted?
The developer used Codex to run 200 fresh sessions, 50 each for a male employee, female employee, male manager, and female manager. The model was asked to give a best estimate percentage of whether the person was slacking off.
What did the statistical analysis show?
Gender, role, and gender-by-role contrasts were significant after Holm correction (all adjusted p < 0.0001). Manager estimates were on average 11.54 percentage points lower than employee estimates.
Which model was tested?
The results came from gpt-5.6-sol with high reasoning effort, on 9/3/2026.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Nvidia CEO Huang Declares AGI Has ArrivedYahoo Finance AI · 1h ago
  • Saudi Arabia's HUMAIN aims to be AI 'Switzerland'Semafor Tech · 1h ago
  • Alibaba releases Qwen-Drive 1.0 driving AI modelTHE DECODER · 1h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleDeepMind agents cheat en masse in math test