Paper Overview
Research Area: NLP Authors: Amir Taubenfeld, Zorik Gekhman, Avigail Grinstein-Dabush, Itay Laish, Ariel Goldstein, Marian Croak, Avinatan Hassidim, Yossi Matias, Amir Feder Published: 2026-09-15 arXiv: 2609.17496
Abstract
LLM assistants are widely used for daily social advice, yet evaluating their social reasoning in such consultation settings remains challenging since:
1. It requires setups where the assistant learns about social situations from subjective user narratives. 2. Social properties, such as others' intentions, typically lack verifiable ground truth.
To address these challenges, the authors introduce Fuse, a multi-agent simulation framework for studying user-mediated social reasoning. In Fuse, a target agent with a hidden motive interacts with other agents, including one representing the user. The user then consults the evaluated assistant to infer the target's motive — providing verifiable ground truth by construction.
Simulation faithfulness is validated through a human study with 24k annotations. The authors applied Fuse to 12 LLMs, demonstrating its analytical utility by systematically isolating key factors:
- User mediation amplifies the inherent difficulty of social reasoning.
- LLMs show systematic sensitivity to biased user framing.
- Models may require more detail than humans to arrive at correct predictions.
- Longer conversations do not always improve performance, even when opportunities for clarifying questions are provided.
--- *Auto-collected on 2026-09-17*