
A researcher corrected a prior LLM resume-screening study after public feedback revealed methodological flaws.
Much of the initial 45 per cent bias flagged was random noise, not genuine bias.
Testing three specific objections across thousands of runs showed extreme baseline instability; reordering prompts and adding blind instructions failed to reduce variance.
What happened
A researcher re-ran a prior study on LLM bias in résumé screening after three commenters challenged the methodology. New experiments with 320, 4,800, and 4,165 runs tested three objections: whether reasoning scores were post-hoc, whether prompt schema order affected stability, and whether placebo prompts drove variance. The researcher found that much of the initial 45 per cent bias flagged three months ago was actually random noise, not genuine bias.
Why it matters
The original study's core claim—that an auditor had flagged 45 per cent of score differences as bias—was based partly on measurement error. When the researcher transplanted positive and negative justifications back into prompts, scores moved 3.62 points in the reasoning's direction 99.7 per cent of the time, confirming that reasoning did influence scores; however, extreme baseline instability revealed that the initial bias count had conflated signal with noise. Testing schema ordering (by moving the score last) and blind instructions failed to reduce variance, and disagreement on hire-versus-no-hire decisions actually rose from 33 per cent to 54 per cent.
What to watch
The study highlights how LLM evaluation can be undermined by instability and measurement artifacts. The researcher's willingness to rerun experiments and revise findings in response to public critique suggests that the field's understanding of LLM bias in hiring decisions may still be incomplete; further methodological refinement may be needed to distinguish real bias from noise.
Ask the AI about this article →
The researcher's original study, published three months ago, concluded that an auditor had flagged 45 per cent of score differences as bias in LLM résumé screening. Public feedback from three commenters—u/kamilc86, u/AssiduousLayabout, and u/hex4def6—challenged the methodology on three specific grounds: that the reasoning justifications might be invented post hoc rather than driving scores; that prompt schema ordering or other structural factors might artificially inflate variance; and that placebo prompts might inflate disagreement. Rather than dismiss these objections, the researcher ran new experiments to test each one.
The first set of experiments, involving 320 runs, showed that when positive and negative justifications were transplanted back into prompts, scores moved 3.62 points in the reasoning's direction 99.7 per cent of the time. This finding confirms that reasoning did influence scores, validating the reasoning-as-cause hypothesis. However, the same experiments revealed extreme baseline instability—meaning that LLM outputs varied wildly even for identical inputs. This instability is crucial: it suggests that much of the 45 per cent bias flagged in the original study was actually noise from this instability, mislabeled as bias.
The second and third experiments tested whether structural changes to the prompts could reduce instability and disagreement. Moving the score last in the prompt (tested across 4,800 runs) had no effect on stability and actually worsened disagreement, raising it from 33 per cent to 54 per cent. Blind prompt instructions also failed. Together, these results indicate that LLM résumé screening is unstable at a fundamental level, and that simple prompt engineering cannot fix it. The researcher's decision to revise and re-release the findings suggests that careful measurement and willingness to correct error are now shaping how this field understands LLM bias.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
Anthropic plans to "match or beat" the size of SpaceX's $75 billion IPO (or $86.2 billion including the over-a…

Pew Research released a study on Thursday finding that over one-third (35%) of English-language web pages publ…

The article argues that non-expert managers and consultants—people whose only exposure to AI comes from ChatGP…

OpenAI's GPT-5.6 Sol, launched July 9, drove a 35 percent revenue increase this quarter, with enterprise reven…

OpenAI is previewing transparent background support for GPT-Image-2 through its API, allowing users to generat…

HP Korea has formed a partnership with Upstage, a large language model (LLM) startup, to advance its localized…
