Paper Overview
- Field: Machine Learning
- Authors: Jagdish Tripathy, Marcus Buckmann
- Published: 2025-05-15
- arXiv: 2505.10888
- Output-internal disconnect: Models show no bias at the output level but retain and amplify demographic representations internally.
- Causal potency: Suppressed bias information is decision-relevant; reinjection at critical layers causes near-complete decision reversals.
- Asymmetry: Interventions strongly affect decisions in one demographic direction but have minimal effect in the reverse.
- Exploitability: The latent bias can be surfaced via adversarial prompt engineering and parameter-efficient fine-tuning.
Abstract (translated summary)
Instruction-tuned language models exhibit behavioral fairness in high-stakes decisions while retaining biased associations in their internal representations. Whether these suppressed representations can affect model outputs—and whether such causal potency is symmetric across demographic groups—remains unknown. This paper investigates open-weight models for mortgage underwriting using matched applications that differ only in racially-associated names, revealing a critical disconnect: models show no output-level bias, yet retain and amplify demographic representations across model layers.
Through activation steering and novel cross-layer interventions, the authors demonstrate that this suppressed information is decision-relevant: when reinjected at critical layers, it produces near-complete decision reversals. Crucially, the latent bias is asymmetric—steering interventions affect decisions in one demographic direction while producing minimal effect in the reverse—and is susceptible to adversarial prompt engineering and parameter-efficient fine-tuning attacks.
Key Findings
Implications
The results show that output-focused behavioral audits are insufficient: fair outputs can mask exploitable internal bias. The authors advocate a two-tier testing framework combining output evaluation with representation analysis for AI governance in high-stakes decisions.
--- *Auto-collected on 2026-05-19.*