Paper Overview
Field: Machine Learning Author: Travis LaCroix Published: 2026-04-22 arXiv: 2604.20805
Abstract
The value alignment problem for artificial intelligence (AI) is often framed as a purely technical or normative challenge, sometimes focused on hypothetical future systems. I argue that the problem is better understood as a structural question about governance: not whether an AI system is aligned in the abstract, but whether it is aligned enough, for whom, and at what cost. Drawing on the principal-agent framework from economics, this paper reconceptualises misalignment as arising along three interacting axes: objectives, information, and principals. The three-axis framework provides a systematic way of diagnosing why misalignment arises in real-world systems and clarifies that alignment cannot be treated as a single technical property of models but an outcome shaped by how objectives are set, how information is distributed, and whose interests are counted in practice.
Key Points
- Reframing alignment as governance: Alignment should be assessed as "aligned enough, for whom, and at what cost" rather than as a binary abstract property of AI systems.
- Three-axis framework from principal-agent theory: Misalignment arises along three interacting axes — objectives, information, and principals — adapted from economics.
- Diagnostic tool: The framework offers a systematic way to diagnose why misalignment emerges in real-world deployed systems.
- Core thesis: Because alignment is pluralistic and context-dependent, and because failures along each axis affect different stakeholders differently, alignment is fundamentally a governance problem rather than a purely technical engineering one.
- Implication: Alignment cannot be "solved" once by technical design; it requires ongoing institutional processes governing how objectives are set, how systems are evaluated, and how affected communities can challenge or reshape those decisions.
*Originally shared on zhichai.net, 2026-04-24.*