
What happened
In over 8,000 trials, GPT-4o chose the face humans rate as more competent 87.83% of the time and the more trustworthy face 72.67%, versus about 62.65% predicted for humans, according to lead author Stephen Lehr of Cangrade.
Why it matters
The AI's preference for faces humans read as competent was markedly stronger than the human baseline, suggesting face-based judgments may be amplified rather than removed when such models are used.
What to watch
The study forced a choice between two faces and used flat 2D stills, so behavior in settings without a direct comparison, or with video or 3D images, is unknown. Whether the bias continues to rise with each new model is the test.
WHO IT HITSRecruiters and hiring teams using AI screening tools, and legal or compliance staff relying on AI-assisted decisions, face a documented risk that these models lean harder on facial appearance than human evaluators do, the study suggests.
Summaries like this, in your inbox every morning.
The study tested whether AI would escape a well-documented human habit: judging people by their faces. Humans tend to read competence or trustworthiness from facial structure even though it does not reflect character, a bias rooted from childhood. Researchers led by Stephen Lehr of Cangrade, an AI recruiting-tool developer, showed AI models pairs of faces tuned to look more or less competent and more or less trustworthy.
The results ran opposite to the hope that AI could remove such bias. GPT-4o leaned toward the faces humans read as competent far more often than humans did themselves. When the team tested monkeys' faces that humans had rated as kind or mean, GPT-4o still picked the kindly rated ones 66% of the time, hinting it had formed a general concept of facial trustworthiness rather than memorizing human faces. Asked which face was more likely to be a serial killer, human trafficker or financial fraudster, GPT-4o chose the less trustworthy-looking face 68.7% of the time, and it picked the more competent-looking face 75.19% of the time in realistic choices such as selecting a university president or a startup to invest in.
The finding that newer models showed more bias, not less, suggests training data alone does not explain it, the team argues — the models appear to amplify it. The stakes hinge on whether this behavior holds outside forced two-face comparisons and on flat images, and for whom: recruiters and legal teams weighing AI tools are the ones most exposed, the authors suggest.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
DeepSeek is bringing more of the software it uses to develop its AI models to Huawei Technologies' Ascend 950…

Nvidia released the Open Agent Safety Platform on September 28, days after CEO Jensen Huang called warnings fr…

OpenAI launched Dots — agents on GPT-6 Astra, each running on its own cloud computer and connecting to 4,000+…

OpenAI released Dots, an always-on agent running autonomously in the cloud on GPT-6 Astra, connecting to over…

OpenAI announced dots, an always-on agent running GPT-6 Astra with its own cloud computer and browser, able to…

Anthropic published a September 29, 2026 review finding Z.ai's downloadable GLM-5.3 built working exploits in…
