English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Fair Outputs, Biased Internals: Causal Potency and Asymmetry of Latent Bias in LLM Mortgage Underwriting

Forum topic · 小凯 · 2026-05-19

Summary

This paper (arXiv:2505.10888) by Jagdish Tripathy and Marcus Buckmann examines whether instruction-tuned language models that appear fair in high-stakes decisions still carry causally potent internal bias. Using matched mortgage applications differing only in racially-associated names, the authors show that open-weight models exhibit no output-level bias yet retain and amplify demographic representations across layers. Through activation steering and novel cross-layer interventions, they demonstrate this suppressed information is decision-relevant: reinjection at critical layers causes near-complete decision reversals. The latent bias is asymmetric—interventions shift decisions strongly in one demographic direction but minimally in the reverse—and is exploitable via adversarial prompt engineering and parameter-efficient fine-tuning. The findings argue that output-focused behavioral audits are insufficient for AI governance in high-stakes domains and motivate a two-tier testing framework combining output evaluation with representation analysis.

Paper Overview

Field: Machine Learning Authors: Jagdish Tripathy, Marcus Buckmann Published: 2025-05-15 arXiv: 2505.10888

Key Findings

  • Instruction-tuned language models show behavioral fairness in high-stakes decisions while retaining biased associations in their internal representations.
  • The authors investigate open-weight models in mortgage underwriting using matched applications that differ only in racially-associated names.
  • A critical disconnect emerges: models show no output-level bias, yet retain and amplify demographic representations across model layers.
  • Through activation steering and novel cross-layer interventions, the suppressed information is shown to be decision-relevant: reinjecting it at critical layers produces near-complete decision reversals.
  • The latent bias is asymmetric: steering interventions affect decisions in one demographic direction while producing minimal effect in the reverse.
  • The suppressed bias is exploitable via adversarial prompt engineering and parameter-efficient fine-tuning.

Implications

These results show that output-focused behavioral audits are insufficient: fair outputs may conceal exploitable internal bias. The authors advocate a two-tier testing framework for AI governance in high-stakes decision-making that combines output-level evaluation with internal representation analysis.

Abstract (Original)

Instruction-tuned language models exhibit behavioural fairness in high-stakes decisions while retaining biased associations in their internal representations. However, whether these suppressed representations can affect model outputs - and whether such causal potency is symmetric across demographic groups - remains unknown. We investigate the use of open-weight models for mortgage underwriting using matched applications that differ only in racially-associated names and reveal a critical disconnect: models show no output-level bias, yet retain and amplify demographic representations across model layers. Through activation steering and novel cross-layer interventions, we demonstrate that this suppressed information is decision-relevant: when reinjected at critical layers, it produces near-co[mplete decision reversals] ...

---

*Auto-collected on 2026-05-19*

Tags

#machine-learning#llm-bias#ai-fairness#arxiv#mortgage-underwriting#interpretability#ai-governance

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620353