LLMs can learn hiring bias themselves, and OpenAI's o3 hit near-max segregation
In a simulated hiring game, researchers found AI stereotypes applicants more than humans, and “fairness” settings barely helped.

Researchers at Princeton University and the University of Chicago tested LLMs, including ChatGPT, Claude, and Gemini, in a simulated hiring scenario. The models rapidly developed demographic segregation, with OpenAI’s reasoning model o3 scoring 1.83 on a 0 to 2 segregation scale.
The unsettling part is not that AI can absorb bias from human history. It is that the system can generate new bias from its own hiring experience, even when every candidate has the same chance of success.
In the study, LLMs were placed into a “mayor of a fictional city” hiring simulation, told to maximize successful hires over 40 rounds across 20 jobs. The researchers say all candidates were equally likely to succeed, yet the models quickly started steering different ethnic groups into different jobs based on early outcomes. That segregation got extreme: on the segregation scale where 2 means every group is confined to its own niche, human participants scored 0.84. The models scored about 65% higher, and OpenAI’s reasoning model o3 scored 1.83, close to the maximum possible.
This matters for a very practical reason. The moment AI starts hiring, the feedback loop changes. In real life, companies do not instantly know whether a hire “worked” after one week. But over time, performance signals trickle in. If the model is then using those signals as training data, it can over-index on whatever patterns it learned early, even if those patterns are accidental or misleading. The paper is effectively showing how quickly a decision system can build a rule of thumb that sounds rational inside the sandbox, but becomes discriminatory once deployed.
So why does it happen? The researchers connect it to a familiar tension in decision-making: the exploration-exploitation dilemma. Humans face it too, like trying a new restaurant versus sticking with your reliable favorite. But in these experiments, LLMs were effectively optimized to generalize from limited information, and that can be a feature in math, coding, and science problems. In social settings, the same instinct becomes the bug. Ryan Liu, a PhD student at Princeton University and a coauthor of the study, said LLMs are “really are eager to create generalizations from limited data,” and that they are “literally” optimized for that.
The study also suggests that higher-reasoning models may intensify the issue. Newer models with higher reasoning capabilities, such as OpenAI’s o3 and DeepSeek’s R1, were described as showing even stronger biases. That does not mean “better reasoning” produces worse behavior in every context. It means that when the system is trained to extract patterns aggressively, it may turn early randomness into confident stereotypes.
And the bias is not just “it treats everyone unfairly.” The models were more likely to stereotype by demographic group than the human participants in the original psychology study. For example, when a model was told an Aima failed as a doctor, a role framed as requiring high levels of warmth and competence, it veered away from hiring all Aimas as doctors. Instead, it began hiring Aimas as janitors, which the model categorized as less warm and competent than doctors. On the segregation scale, that kind of drift is exactly what pushes the system toward niche confinement: the model learns that different groups belong in different boxes, even though success rates were equal.
One of the most relevant operational details is what does not work. The study found that telling the model to be fair did not change its behavior much. As Liu put it, “Either it can’t put these values into action or that process is being submerged under the tendency to try to optimize for the goal of getting the most correct hires.” Translation for busy decision-makers: fairness instructions are not magic spells. If the optimization objective and the training dynamics reward “maximize successful hires,” the model may treat fairness as noise or as something it can never effectively implement.
What did work was more structural. Promising an additional bonus for diverse hiring made the models far less biased. The researchers frame the lesson as designing goals that “incorporate desirable social values in order to make the large language model act in socially desirable ways,” and that is a different approach than simply adding a statement about fairness. The models also became less biased when given more personal information relevant to an individual’s ability to adapt in a separate experiment about resettlement across Canada. When provided relevant attributes like age and education, the models were less likely to segregate by ethnicity. When given irrelevant attributes like hair color and tattoo shape, they largely fell back to sorting by ethnicity.
This is landing right as AI hiring tools become more common and more capable. The source notes that chatbots are gaining improved memory and personalization features, which creates new ways biases can accumulate. Angelina Wang, a computer scientist at Cornell University who did not work on the study, said when chatbots draw on previous conversation history, they can “over-index on the same kinds of behaviors it’s experienced before” and form biases. And she added a hard constraint that matters for product teams: users want chatbots to remember what they say, so “simply having chatbots remember less isn’t a fix,” because the question becomes finding the “right amount” of memory.
From a regulatory perspective and board-level risk standpoint, the takeaway is sharper than “be careful with biased training data.” The study points to “novel biases,” meaning biases that no human ever explicitly taught the system could become “ever present” once the AI learns from experience. As Wang put it, the finding is a “really serious implication that they should grapple with,” especially as LLMs are deployed to screen résumés and even conduct interviews. The system can form and reinforce discriminatory patterns without any direct “bias label” being present, simply because it is optimizing over time.
OpenAI and Anthropic did not respond to requests for comment, but the underlying mechanism is the same across the tools: a model that learns from outcomes can misinterpret early signals. For executives and boards, this shifts the conversation. You are not just auditing static model outputs. You are auditing feedback loops, data collection pipelines, and the incentives embedded in hiring objectives. If you are rolling out AI systems that decide who gets hired, who gets a loan, or who gets parole, this study is a warning about where the bias may come from next: not from your training set, but from your system’s own learning.
This story's Key Insights and Take-aways are locked.
Create a free account to unlock Executive Actions for one credit.
Register to UnlockAlways free for Executives Club members. Join the Club
More in Business

Uber buys Delivery Hero for nearly $15B, vaulting to top food delivery outside China
The deal doubles Uber's dual-services footprint and pushes a ride-and-eats bundling play into 50 more markets.

Epic and Google drop settlement bid, forcing rival Android app stores by July 22
Google told the court it is ready to carry third-party app stores starting Wednesday, July 22.

SK Hynix opens at $170, raises $26.5B, and tops foreign IPO records
In Friday's Wall Street debut, SK Hynix turns AI RAM demand into a $26.5B fundraising moment that rewrites comps.

