Paper Overview
- Field: Machine Learning
- Authors: Kevin Kingslin, Anish Natekar, Ashutosh Ranjan
- Published: 2026-06-26
- arXiv: 2606.28294
- Gathers multiple competing rationales through structured persona debate
- Provides a broader and more expressive account of the factors influencing each preference comparison
- Derives clearer and more comprehensive steering principles from these richer signals
- Uses the principles to guide decision modeling with LLM-based and decision-tree judges
Background
Preference-based alignment often struggles to capture the reasoning that underlies human judgments. Many evaluations rely on multiple interacting criteria, yet pairwise labels reveal only the final choice rather than the considerations that shape preferences. Inverse Constitutional AI (ICAI) improves interpretability by summarizing preferences into natural-language principles, but its single-pass explanations miss much of the nuance involved in complex decisions.
Method: Democratic ICAI
The paper introduces Democratic ICAI, a novel approach that:
Results
Experiments on the creative preference benchmarks MuCE-Pref and LiTBench, spanning multiple creative task categories, show that Democratic ICAI produces more faithful preference structures. Compared with deliberative-prompting and principle-based baselines, it improves average preference prediction across tasks while producing constitutions that LLM annotators prefer.
---
*Auto-collected on 2026-06-30.*