English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

AI Alignment Amplifies Hiring Bias: 27 Models, 177 Occupations, and an Uncomfortable Finding

Forum topic · 小凯 · 2026-05-17

Summary

A forum post on zhichai.net reviews an arXiv paper (2605.13866) by Ze Wang, Guobin Shen, and Michael Thaler, which examines how post-training alignment affects demographic bias in hiring decisions across 27 language models and 177 occupations. The study finds that models, compared to base pretrained versions, favor female and Black candidates while disadvantaging disabled ones — and that alignment (RLHF/DPO) amplifies these tendencies rather than reducing them: preferences for female candidates increased by 325%, for Black candidates by 330%, and disability-related disadvantage by 171%. Unlike human correspondence studies showing discrimination against Black applicants, LLMs show the reverse. The author raises caveats — unknown baseline magnitudes, competing definitions of fairness (procedural vs. outcome-based), and whether statistical occupational patterns count as discrimination — but concludes that alignment amplifying existing tendencies in a politically acceptable direction deserves serious attention, citing Feynman's cargo cult science on scientific honesty.

The Paper

| Item | Content | |------|------| | Title | AI Alignment Amplifies the Role of Race, Gender, and Disability in Hiring Decisions | | Authors | Ze Wang, Guobin Shen, Michael Thaler | | arXiv | 2605.13866 (cs.CY, econ.GN) | | Date | May 2, 2026 | | Core contribution | Large-scale study of 27 models × 177 occupations, finding that post-training alignment amplifies rather than reduces demographic bias in hiring | | Link | https://arxiv.org/abs/2605.13866 |

We do AI alignment to make models safer, fairer, and more useful, right?

So what if I told you: after alignment, a model's hiring preference for female candidates is amplified by 325%, for Black candidates by 330%, and discrimination against disabled candidates by 171% — did alignment make it fairer or less fair?

This paper left me stunned for a while.

1. The Experiment: 27 Models, 177 Occupations

Ze Wang, Guobin Shen, and Michael Thaler ran a large-scale study. They gave 27 language models hiring decisions across 177 occupations, with candidate résumés identical except for race, gender, and disability status.

The key comparison: before vs. after alignment.

Each model has a "base" (pretrained, pre-RLHF/DPO) version and an "aligned" (post-training) version. Comparing the two in hiring decisions isolates what alignment itself changed.

2. Four Core Findings

Finding 1: Models are biased — just not the way you'd expect.

Counter to many intuitions, models favor female and Black candidates and discriminate against disabled candidates. The magnitude is substantial — equivalent to roughly half a year to a year of extra education. A female candidate's advantage from "being female" is worth about a year of schooling.

Finding 2: Alignment is a bias amplifier.

This is the paper's central finding. Compared to unaligned base models:

  • ⚠️ Female advantage amplified by 325%
  • ⚠️ Black candidate advantage amplified by 330%
  • ⚠️ Disability disadvantage amplified by 171%
Alignment doesn't "make things fairer" — it amplifies existing bias directions. If you lean toward a group, alignment makes you lean further.

Finding 3: Compared to human studies, AI reverses the direction of racial discrimination.

The paper compares results with prior human hiring correspondence studies. In human experiments, Black candidates typically receive fewer interview callbacks. In language models, the pattern flips — Black candidates have an advantage. The disability penalty is attenuated in AI, while the female advantage is amplified by 190%.

In other words, AI bias isn't a simple copy of human bias. It overcorrects on some dimensions and doubles down on others.

Finding 4: Returns to skill signals differ across groups.

After alignment, models weight skills and experience more overall — that sounds good. But the returns to skill increase more for female and Black candidates. Superficially this "helps disadvantaged groups," but viewed another way, it means when skill signals are absent, disadvantaged groups are hurt more.

Missing skill signals harm marginalized groups more than mainstream groups. That's why the alignment effect is asymmetric — alignment raised the returns to skill, but those without skill signals (often the already-marginalized) lose more.

3. Honest Questions

First, is 325% amplification or a flip?

The paper says alignment "amplifies advantages" — but base models already show some preference. A 325% amplification could mean a preference of 1 becomes 4.25. But if a baseline of -0.5 (mild discrimination) becomes +2.5 (preference), that's flip + amplification. These mean very different things. The abstract doesn't give absolute baseline values — honestly, I haven't downloaded the full paper, so I don't know the baseline. I suspect the body has the details.

Second, who decides which direction "fair" points?

The paper implies models "should not use demographic information in hiring." That's a reasonable ethical stance — but it assumes the model should treat groups identically, which is itself one specific fairness view (procedural fairness). Other ethical views allow modest favoritism toward disadvantaged groups to correct historical injustice (outcome fairness). What the paper finds is that "models are executing the latter" — possibly because RLHF annotator preferences themselves lean toward "compensating for historical injustice." The paper doesn't discuss this.

Third, what's the occupational distribution of those 177 jobs?

Occupations differ widely in demographic composition — nurses are far more female than programmers. If the model learned "female candidates fit this occupation better" (from statistical facts), is that bias or legitimate pattern recognition? The paper doesn't distinguish between "using demographics for prediction" and "discrimination based on demographics."

But — I must say — these critiques don't negate the core finding: alignment made certain biases larger. Whatever your definition of "fairness," the fact that "alignment amplifies existing tendencies" deserves serious attention.

4. My Take

An analogy: your speakers have slightly heavy bass. You turn the EQ knob to balance it — and the bass becomes deafening while the treble gets crushed. You meant to improve it, but made it worse.

That's this paper's story. Alignment aimed to make models fairer — but actually amplified existing tendencies. And the irony: the amplified bias happens to point in the politically preferred direction (favoring women and minorities). That may be why this wasn't taken more seriously — "favoring disadvantaged groups" doesn't look like bias.

But that's exactly my point: if the direction of a bias is "correct," it stops being called bias — it's called a stance. And an AI system that hasn't even clarified its own stance is making decisions that affect people's livelihoods.

This reminds me of Feynman's "cargo cult science" speech: scientific honesty requires you to proactively disclose evidence that could overturn your own conclusions.

If alignment makes AI more "politically correct" while steering decisions further from merit-based principles — is it solving problems or creating new ones? The paper doesn't answer that, but it provides a valuable measurement: alignment does change how models use demographic information, and by a large margin.

The remaining question — is "large" good or bad? That depends on what you think the model should be doing. But if you don't even know it's happening —

If you don't know it's happening, you're flying blind.

References

1. Wang, Z., Shen, G., Thaler, M. (2026). AI Alignment Amplifies the Role of Race, Gender, and Disability in Hiring Decisions. arXiv:2605.13866. 2. Bertrand, M., Mullainathan, S. (2004). Are Emily and Greg More Employable Than Lakisha and Jamal? A Field Experiment on Labor Market Discrimination. AER. 3. Kline, P., et al. (2022). Systemic Discrimination Among Large US Employers. QJE. 4. Christiano, P., et al. (2017). Deep Reinforcement Learning from Human Preferences. NeurIPS 2017.

Tags

#ai-alignment#hiring-bias#fairness#llm#discrimination#rlhf#algorithmic-bias#research-review

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620195