Paper Overview
Field: Natural Language Processing (NLP) Authors: Ilana Nguyen, Harini Suresh, Thema Monroe-White Published: 2025-04-28 arXiv: 2504.19772
Abstract
Large language models (LLMs) are increasingly used for text generation tasks ranging from everyday applications to high-stakes enterprise and government use cases, including simulated interviews with asylum seekers. While many works highlight the new potential applications of LLMs, there are risks of LLMs encoding and perpetuating harmful biases about non-dominant communities across the globe.
The authors' findings demonstrate the presence of persistent representational harms by national origin, including harmful stereotypes, erasure, and one-dimensional portrayals of Global Majority identities. Minoritized national identities are simultaneously underrepresented in power-neutral narratives and overrepresented in subordinated character portrayals, which are over fifty times more likely to appear than dominant portrayals.
When U.S. nationality cues are present in input prompts, the degree of harm is amplified. Notably, these harms cannot be explained by a sycophancy effect: even when U.S. nationality cues in prompts are replaced with non-U.S. national identities, the U.S.-centric bias persists.
Key Points
- Scope of use: LLMs are increasingly deployed in sensitive government and enterprise workflows, such as simulating asylum-seeker interviews, where biased outputs have real-world consequences.
- Representational harms identified: The paper documents three categories of harm toward Global Majority identities:
- Harmful stereotypes
- Erasure (underrepresentation in neutral contexts)
- One-dimensional portrayals
- Asymmetric portrayal: Minoritized national identities appear in subordinated roles more than 50× as often as in dominant roles, while also being underrepresented in power-neutral narratives.
- Prompt sensitivity: Introducing U.S. nationality cues in prompts amplifies representational harms.
- Beyond sycophancy: The U.S.-centric bias persists even after replacing U.S. cues with non-U.S. national identities, indicating the bias is not merely a reflection of user prompting but is encoded in the models themselves.
- Implication: Standard sycophancy-mitigation strategies are insufficient; targeted interventions are needed to address structural representational harms in LLM-generated narratives.