Paper Overview
- Field: Machine Learning
- Authors: Zixiang Xu, Sixian Li, Huaxing Liu, Xiang Wang, Shuai Li
- Published: 2026-07-13
- arXiv: 2607.11871
- Baseline judging inputs occupy a compact activation manifold.
- Biased inputs are displaced along a low-dimensional, bias-type-specific subspace.
- This subspace sharpens with model depth and is consistently recovered by three families of estimators.
- Steering hidden states along this subspace drives scoring in both directions.
- Forward shifts reproduce biased scores on clean inputs.
- Reverse shifts restore baseline scores on biased inputs.
- Matched random directions produce shifts an order of magnitude smaller.
- A simple linear projection along the same bias-direction features can predict judge failures on three fully unseen benchmarks.
- It significantly outperforms text-based alternatives.
Abstract (translated)
Existing studies of LLM-as-judge scoring bias work predominantly at the input-output level: they perturb inputs, measure score deltas, and propose prompt-level mitigations. This paper argues that the same biases admit a representation-level account in the judge's hidden state — complementary to the input-output view and operationally more useful. Three findings are reported across seven judge models, seven bias types, and nine benchmarks.
Key Findings
1. Geometry
2. Causal Control
3. Operational Utility
Conclusion
Reading bias as activation geometry rather than input-output noise unifies geometric structure, causal control, and predictive utility within a single mechanistic framework.
---
*Auto-collected on 2026-07-15*