Participatory Moral AI Is Not Neutral: The Invisible Hand of Developer Choices
Research area: ML Authors: Taenyun Kim, Edyta Bogucka, Daniele Quercia arXiv: 2508.08540
Overview
As AI systems make more morally loaded decisions across society, one response has been moral preference elicitation: researchers poll participants on hypothetical dilemmas and use the aggregated votes to train a policy that an AI model then applies at scale.
However, before any vote is cast, developers make three key choices in the moral AI elicitation pipeline:
- Feature scoping — which features go to a vote
- Voter sampling — which voters to include
- Question framing — how the question is presented
Findings
The authors examined each of these choices in a two-stage study (N=809) covering three deployment scenarios: AI kidney allocation, AI agents simulating absent workers, and generative AI depictions of deceased people.
1. Feature scoping: Morally relevant features shift across scenarios, suggesting that feature patterns should not be assumed to transfer across deployment domains. 2. Voter sampling: Roughly one-third of feature preferences vary with political ideology, and some differences even reverse direction. The ideological composition of the voter pool therefore shapes the final aggregated preference profile. 3. Question framing: The wording of elicitation questions can shrink or widen ideological gaps by up to a full scale point, and framing conditions also change how moral foundations relate to participants' judgments.
Conclusion
Taken together, these findings show that vote-based alignment cannot achieve fair or transparent AI through aggregation alone. At minimum, every stage of the moral AI elicitation pipeline should be audited and disclosed.
Source: arXiv:2508.08540