LLMs may develop hiring bias from experience, not just training data
MIT Tech Review AI · rss · 2026-07-20
A new study suggests that LLMs may be more likely than humans to form hiring biases.
Researchers at Princeton and the University of Chicago ran models such as ChatGPT, Claude, and Gemini through a simulated hiring game based on a psychology experiment. The models were asked to maximize successful hires across 40 rounds, but all candidates were equally likely to succeed. Despite that, the models quickly started to segregate candidates by fictional ethnic group into different job niches after limited experience.
Key findings:
- The models showed stronger stereotyping than humans from the original study; on the segregation scale, human participants scored 0.84, while the models were about 65% higher, with OpenAI’s o3 at 1.83, close to the maximum.
- Newer reasoning models, including o3 and DeepSeek R1, showed even stronger bias.
- Simply telling the models to be fair had little effect.
- Adding a bonus for diverse hiring reduced bias a lot.
- Giving the models more relevant personal information about candidates also reduced stereotyping, while irrelevant details did not.
The article argues this matters for real-world uses like résumé screening, hiring, loans, and parole: as systems gain memory and experience, they may learn biased shortcuts on their own, not just inherit human prejudice.
More from Research
- Structural ensembles beat single predictions in TCR:pMHC generalization study — quaidmorris · 2026-07-22
- RSS launches under OMSF to push structural biology data modeling at scale — MoAlQuraishi · 2026-07-22
- enFoldX tops 8 neoantigen scans and an unseen-peptide benchmark — quaidmorris · 2026-07-22
- enFoldX reaches AUC 0.82 on human VDJdb and transfers to mouse at 0.76 — quaidmorris · 2026-07-22
- enFoldX gains accuracy as AF3 ensemble disagreement rises for non-binders — quaidmorris · 2026-07-22
- A 3D ray plot shows how hard this Jacobian counterexample is to read — moultano · 2026-07-22