English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Inside the Unfair Judge: Mechanistic Interpretability of LLM-as-Judge Scoring Bias

Forum topic · 小凯 · 2026-07-15

Summary

This paper, Inside the Unfair Judge (arXiv:2607.11871), argues that LLM-as-judge scoring biases admit a representation-level explanation in the judge's hidden states, complementing conventional input-output perturbation studies. Across seven judge models, seven bias types, and nine benchmarks, the authors report three findings. First, geometry: baseline judging inputs occupy a tight activation manifold, while biased inputs are displaced along a low-dimensional, type-specific subspace that sharpens with model depth and is consistently recovered by three families of estimators. Second, causal control: steering hidden states along this subspace drives scoring bidirectionally—forward shifts reproduce biased scores on clean inputs, reverse shifts restore baseline scores on biased inputs, and matched random directions produce effects an order of magnitude smaller. Third, operational utility: simple linear probes trained on these bias directions can predict judge failures on three fully unseen benchmarks, substantially outperforming text-based alternatives. The work unifies geometric structure, causal control, and predictive utility in a single mechanistic framework for judging bias.

Paper Overview

  • Field: Machine Learning
  • Authors: Zixiang Xu, Sixian Li, Huaxing Liu, Xiang Wang, Shuai Li
  • Published: 2026-07-13
  • arXiv: 2607.11871
  • Abstract (translated)

    Existing studies of LLM-as-judge scoring bias work predominantly at the input-output level: they perturb inputs, measure score deltas, and propose prompt-level mitigations. This paper argues that the same biases admit a representation-level account in the judge's hidden state — complementary to the input-output view and operationally more useful. Three findings are reported across seven judge models, seven bias types, and nine benchmarks.

    Key Findings

    1. Geometry

  • Baseline judging inputs occupy a compact activation manifold.
  • Biased inputs are displaced along a low-dimensional, bias-type-specific subspace.
  • This subspace sharpens with model depth and is consistently recovered by three families of estimators.
  • 2. Causal Control

  • Steering hidden states along this subspace drives scoring in both directions.
  • Forward shifts reproduce biased scores on clean inputs.
  • Reverse shifts restore baseline scores on biased inputs.
  • Matched random directions produce shifts an order of magnitude smaller.
  • 3. Operational Utility

  • A simple linear projection along the same bias-direction features can predict judge failures on three fully unseen benchmarks.
  • It significantly outperforms text-based alternatives.

Conclusion

Reading bias as activation geometry rather than input-output noise unifies geometric structure, causal control, and predictive utility within a single mechanistic framework.

---

*Auto-collected on 2026-07-15*

Tags

#llm-as-judge#mechanistic-interpretability#activation-steering#scoring-bias#representation-analysis#machine-learning#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178395146