AIToday

AI hiring screeners stereotype job candidates more than humans do

MIT Technology Review AI16h ago
AI hiring screeners stereotype job candidates more than humans do

Key takeaway

A study by Princeton and University of Chicago researchers found that large language models screening job candidates form biases more readily than humans, segregating ethnic groups into specific job types based on limited early experience. OpenAI's o3 reasoning model showed particularly strong stereotyping, scoring 1.83 on a segregation scale where 2 represents complete occupational segregation by ethnicity. The finding is urgent as companies increasingly use AI for résumé screening and interviews—and as chatbots gain memory features that let them learn and reinforce patterns over time.

Summaries like this, in your inbox every morning.

Sign up free →

3 Key Points

  • What happened

    Researchers at Princeton University and the University of Chicago tested LLMs including ChatGPT, Claude, and Gemini in a simulated hiring game where models screened candidates from fictional ethnic groups for 20 different jobs over 40 rounds. All candidates were equally likely to succeed, but the models quickly learned to segregate ethnic groups into specific jobs—for example, steering one group away from doctor positions after a single failure and toward janitor roles instead.

  • Why it matters

    The models showed roughly 65% higher segregation than human participants in the original psychology study, with OpenAI's reasoning model o3 scoring 1.83 on a segregation scale where 2 represents complete confinement. LLMs are trained to generalize from limited data—a strength for logic puzzles but a liability in hiring, where early patterns can harden into unfair stereotypes. As companies increasingly deploy AI to screen résumés and conduct interviews, these learned biases could affect real job applicants without humans ever teaching the model to discriminate.

  • What to watch

    The research identified two levers that reduced bias: promising models a bonus for diverse hiring made them far less biased, and providing relevant personal information about candidates (age, education) rather than irrelevant details (hair color, tattoos) decreased ethnic segregation. The study was published at ICML in Seoul in July; real-world impact remains uncertain because AI screeners don't get instant feedback on hiring success the way the experiment did.

In Depth

A team of researchers from Princeton University and the University of Chicago conducted an experiment to test whether large language models form biases in hiring decisions. They adapted a psychology study originally designed to explore how humans develop stereotypes and applied it to three major LLMs—ChatGPT, Claude, and Gemini—as well as newer reasoning models OpenAI's o3 and DeepSeek's R1.

In the experiment, each model was told it had been hired as a consultant by the mayor of a fictional city and was asked to help hire people for 20 jobs, including doctors, lawyers, child-care aides, and janitors. Candidates came from four fictional ethnic groups: Tufa, Aima, Reku, and Weki. In each of 40 rounds, a new job opening appeared with four candidates, one from each group. After the model hired a candidate, it learned whether that person succeeded at the job and then moved to the next round. Crucially, all candidates were equally likely to succeed at every job—a fact the models did not know.

The results were stark. The models quickly began segregating candidates from different groups into different jobs based on early observations. For instance, when a model learned that an Aima candidate had failed as a doctor (a role the model classified as requiring high warmth and competence), it stopped hiring Aimas for doctor positions and instead began steering them toward janitor roles, which it classified as lower in those qualities. On a segregation scale where 2 represents complete occupational confinement by ethnicity, human participants in the original study scored 0.84. The LLMs scored roughly 65% higher: OpenAI's o3 scored 1.83, close to the maximum possible. The newer, more capable models showed even stronger biases than earlier versions.

Ryan Liu, a PhD student at Princeton and coauthor of the study, explained that LLMs are "really eager to create generalizations from limited data" because "that's literally a lot of what they're optimized for." LLMs are trained on mathematics, coding, and science problems—domains where generalizing from a few examples is a reliable strategy. This same instinct that helps them solve logic puzzles makes them quick to stereotype in social contexts. Liu emphasized: "When LLMs rush to generalize in social settings, that's when things tend to go wrong."

The researchers then tested interventions. Simply telling the models to be fair made little difference. Liu noted: "Either it can't put these values into action or that process is being submerged under the tendency to try to optimize for the goal of getting the most correct hires." However, promising the models an additional bonus for diverse hiring made them far less biased. In a separate experiment, when models were given relevant personal information about candidates—such as age and education—they were less likely to segregate by ethnicity. When given irrelevant information, such as hair color and tattoo shape, they largely reverted to ethnic sorting.

The study, published at ICML in Seoul in July, carries immediate real-world implications. As companies increasingly deploy LLMs to screen résumés and conduct interviews, the finding that models develop biases from their hiring experience "is a really serious implication that they should grapple with," according to Angelina Wang, a computer scientist at Cornell University who did not work on the study. Wang flagged an additional concern: as chatbots gain improved memory and personalization features, they can "over-index on the same kinds of behaviors it's experienced before" and form biases. Yet removing memory is not a practical solution, as "users want chatbots to remember what they say."

A key uncertainty remains how these findings translate to real hiring. In the experiment, models received immediate feedback on whether hires succeeded. In practice, companies may take a long time to determine whether a new employee succeeds, and feedback can be incomplete or skewed. Nevertheless, when feedback does arrive, a model could still misinterpret limited results and carry those misreadings forward. As Liu concludes: "These novel biases—they're sort of ever present."

Context & Analysis

The research addresses a critical gap in how we understand AI bias. While AI companies and researchers have long known that LLMs absorb human biases from training data, this study reveals that models can independently develop stereotypes through learned experience—even when training data contains no explicit bias. The hiring simulation exploited a core feature of how LLMs work: they are optimized to generalize from limited examples, a strength in domains like mathematics and coding where sparse examples are reliable signals. But in social judgment, the same generalization instinct becomes a liability. Ryan Liu, a coauthor, notes that LLMs are "really eager to create generalizations from limited data," and once a model observes a single failure (an Aima candidate failing as a doctor), it extrapolates that pattern into sweeping avoidance, steering the entire group toward lower-status roles.

The severity of this bias—65% higher segregation than human participants—was unexpected. Newer reasoning models like o3 and DeepSeek's R1, which excel at logic and planning, exhibited the strongest biases, suggesting that advanced reasoning capability does not naturally protect against social stereotyping. Importantly, the study also tested interventions: simply instructing models to "be fair" had little effect, implying that fairness values do not override the optimization pressure for accurate predictions. However, reframing the goal (bonus for diverse hires) or providing richer, more relevant context (education and age rather than tattoos) both reduced bias sharply. This indicates that bias is not baked into the model weights but flows from the decision architecture and feedback signal.

The real-world stakes are substantial. Unlike the lab, an AI screener deployed in hiring does not receive instant feedback on whether a candidate succeeds. Feedback is delayed and often incomplete, yet when it arrives, the model may still overinterpret limited signals. As chatbots gain persistent memory and personalization, Angelina Wang, a Cornell computer scientist, warns they will "over-index on the same kinds of behaviors experienced before." The challenge is acute: users want systems to remember prior conversations, yet memory is the mechanism through which learned biases accumulate. The tension between utility and fairness remains unresolved.

FAQ

What models were tested in this study?
Researchers tested ChatGPT, Claude, Gemini, OpenAI's reasoning model o3, and DeepSeek's R1. Newer models with higher reasoning capabilities, such as o3 and R1, showed even stronger biases than earlier versions.
What made the AI models less biased in the experiment?
Promising the models an additional bonus for diverse hiring made them far less biased. The models also became less biased when given relevant personal information about candidates, such as age and education.
How did the AI models' bias compare to humans?
On the study's segregation scale where 2 means every group has been completely confined to its own job niche, human participants scored 0.84. The models scored roughly 65% higher, with OpenAI's o3 scoring 1.83, close to the maximum possible.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Discussion

No comments yet. Be the first to share your thoughts!

Log in to join the discussion

Related Articles

Stay ahead with AI news

Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.

Get Started Free

Free · takes 30 seconds · unsubscribe anytime

1 minute a day. The AI essentials.

200+ sources · Email / LINE / Slack

Get it free →