Paper Overview
Field: AI/ML Authors: Yuhe Wu, Guangyu Wang, Yujie Chen Published: 2026-09-06 arXiv: 2509.00006
Summary
People increasingly turn to large language models (LLMs) for everyday advice, making ethically charged interpersonal problems a practical moral-advisory context. Most prior work has studied this setting through single-turn judgments or pressure-laden rebuttals—assumptions that poorly match how guidance is sought in the real world. This leaves open a key question: can narration alone, without any explicit opposing position, shift a model's judgment during multi-turn moral consultation?
In real-world moral-conflict conversations, one party typically offers a self-justifying account that unfolds over multiple turns, creating information asymmetry. The paper introduces narrative captivity, a failure mode in which a model treats an unopposed one-sided account as complete and aligns with the narrator's interpretation instead of seeking missing perspectives.
Key Findings
- The authors construct a benchmark of 5,078 interpersonal conflict scenarios covering six moral dimensions.
- Across 17 LLMs, narrative captivity is pervasive: final-state judgments under multi-turn narration shift by 25 percentage points on average relative to matched single-turn baselines.
- Stage-level analysis identifies preference optimization as the primary contributing factor.
- Four inference-time strategies offer only partial mitigation.
Takeaway
The authors hope this project will promote LLM advisors that maintain independent judgment in real-world consulting scenarios, rather than being carried along by a single party's narrative.
--- *Auto-collected on 2026-09-06*