Paper Overview
Field: Machine Learning Authors: Jagdish Tripathy, Marcus Buckmann Published: 2025-05-15 arXiv: 2505.10888
Key Findings
- Instruction-tuned language models show behavioral fairness in high-stakes decisions while retaining biased associations in their internal representations.
- The authors investigate open-weight models in mortgage underwriting using matched applications that differ only in racially-associated names.
- A critical disconnect emerges: models show no output-level bias, yet retain and amplify demographic representations across model layers.
- Through activation steering and novel cross-layer interventions, the suppressed information is shown to be decision-relevant: reinjecting it at critical layers produces near-complete decision reversals.
- The latent bias is asymmetric: steering interventions affect decisions in one demographic direction while producing minimal effect in the reverse.
- The suppressed bias is exploitable via adversarial prompt engineering and parameter-efficient fine-tuning.
Implications
These results show that output-focused behavioral audits are insufficient: fair outputs may conceal exploitable internal bias. The authors advocate a two-tier testing framework for AI governance in high-stakes decision-making that combines output-level evaluation with internal representation analysis.
Abstract (Original)
Instruction-tuned language models exhibit behavioural fairness in high-stakes decisions while retaining biased associations in their internal representations. However, whether these suppressed representations can affect model outputs - and whether such causal potency is symmetric across demographic groups - remains unknown. We investigate the use of open-weight models for mortgage underwriting using matched applications that differ only in racially-associated names and reveal a critical disconnect: models show no output-level bias, yet retain and amplify demographic representations across model layers. Through activation steering and novel cross-layer interventions, we demonstrate that this suppressed information is decision-relevant: when reinjected at critical layers, it produces near-co[mplete decision reversals] ...
---
*Auto-collected on 2026-05-19*