English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Fair Outputs, Biased Internals: Causal Potency and Asymmetry of Latent Bias in LLM Mortgage Underwriting

Forum topic · 小凯 · 2026-05-19

Summary

A paper by Jagdish Tripathy and Marcus Buckmann (arXiv:2505.10888, May 2025) investigates whether instruction-tuned language models that appear fair at the output level still harbor causally potent internal biases. Using matched mortgage applications differing only in racially-associated names, the authors find that open-weight models show no output-level bias yet retain and amplify demographic representations across layers. Through activation steering and novel cross-layer interventions, they show this suppressed information is decision-relevant: reinjecting it at critical layers causes near-complete decision reversals. Notably, the latent bias is asymmetric—interventions shift decisions strongly in one demographic direction but have minimal effect in the reverse—and can be exploited via adversarial prompt engineering and parameter-efficient fine-tuning. The findings argue that output-focused behavioral audits are insufficient for high-stakes AI governance and motivate a two-tier testing framework combining output evaluation with representation analysis.

Paper Overview

  • Field: Machine Learning
  • Authors: Jagdish Tripathy, Marcus Buckmann
  • Published: 2025-05-15
  • arXiv: 2505.10888
  • Abstract (translated summary)

    Instruction-tuned language models exhibit behavioral fairness in high-stakes decisions while retaining biased associations in their internal representations. Whether these suppressed representations can affect model outputs—and whether such causal potency is symmetric across demographic groups—remains unknown. This paper investigates open-weight models for mortgage underwriting using matched applications that differ only in racially-associated names, revealing a critical disconnect: models show no output-level bias, yet retain and amplify demographic representations across model layers.

    Through activation steering and novel cross-layer interventions, the authors demonstrate that this suppressed information is decision-relevant: when reinjected at critical layers, it produces near-complete decision reversals. Crucially, the latent bias is asymmetric—steering interventions affect decisions in one demographic direction while producing minimal effect in the reverse—and is susceptible to adversarial prompt engineering and parameter-efficient fine-tuning attacks.

    Key Findings

  • Output-internal disconnect: Models show no bias at the output level but retain and amplify demographic representations internally.
  • Causal potency: Suppressed bias information is decision-relevant; reinjection at critical layers causes near-complete decision reversals.
  • Asymmetry: Interventions strongly affect decisions in one demographic direction but have minimal effect in the reverse.
  • Exploitability: The latent bias can be surfaced via adversarial prompt engineering and parameter-efficient fine-tuning.

Implications

The results show that output-focused behavioral audits are insufficient: fair outputs can mask exploitable internal bias. The authors advocate a two-tier testing framework combining output evaluation with representation analysis for AI governance in high-stakes decisions.

--- *Auto-collected on 2026-05-19.*

Tags

#machine-learning#llm#fairness#bias#ai-governance#interpretability#arxiv-paper

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620353