Paper Overview
Field: AI/ML Authors: Yuhe Wu, Guangyu Wang, Yujie Chen Published: 2026-09-06 arXiv: 2509.00006
Summary
People increasingly turn to large language models (LLMs) for everyday advice, making ethically charged interpersonal problems a practical moral-advisory context. Most prior work has studied this context through single-turn judgments or pressure-laden rebuttals—assumptions that poorly match how guidance is sought in real-world contexts. These assumptions leave unclear whether narration alone, without an explicit opposing position, can shift model judgments during multi-turn moral consultation. Yet real-world moral-conflict conversations often elicit one party's self-justifying account, which can unfold over multiple turns and create information asymmetry.
The authors introduce narrative captivity, a failure mode in which a model treats an unopposed one-sided account as complete and aligns with the narrator's interpretation rather than seeking the missing perspective.
Key Findings
- Benchmark: 5,078 interpersonal conflict scenarios covering six moral dimensions.
- Prevalence: Across 17 LLMs, narrative captivity is widespread—final-state judgments under multi-turn narration shift by an average of 25 percentage points versus matched single-turn baselines.
- Cause: Stage-level analysis identifies preference optimization as the primary contributing factor.
- Mitigation: Four inference-time strategies provide only partial relief.
--- *Auto-collected on 2026-09-06*