English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Verifiable Social Reasoning for LLM Assistants: The Fuse Multi-Agent Simulation Framework

Forum topic · 小凯 · 2026-09-17

Summary

Researchers from Google and academic collaborators introduce Fuse, a multi-agent simulation framework for evaluating how LLM assistants perform social reasoning in user consultation settings. In Fuse, a target agent with a hidden motive interacts with other agents, including one representing the user; the user then consults the assistant under evaluation to infer the target's motive, providing verifiable ground truth by construction. Simulation faithfulness was validated via a human study with 24,000 annotations. Experiments across 12 LLMs show that user-mediated consultation amplifies the inherent difficulty of social reasoning, models are systematically sensitive to biased user framing, may need more detail than humans to reach correct predictions, and longer conversations do not always improve performance even when clarifying questions are possible. The framework and a 21,000-example dataset are open-sourced (arXiv:2609.17496).

Paper Overview

Research Area: NLP Authors: Amir Taubenfeld, Zorik Gekhman, Avigail Grinstein-Dabush, Itay Laish, Ariel Goldstein, Marian Croak, Avinatan Hassidim, Yossi Matias, Amir Feder Published: 2026-09-15 arXiv: 2609.17496

Abstract

LLM assistants are widely used for daily social advice, yet evaluating their social reasoning in such consultation settings remains challenging since:

1. It requires setups where the assistant learns about social situations from subjective user narratives. 2. Social properties, such as others' intentions, typically lack verifiable ground truth.

To address these challenges, the authors introduce Fuse, a multi-agent simulation framework for studying user-mediated social reasoning. In Fuse, a target agent with a hidden motive interacts with other agents, including one representing the user. The user then consults the evaluated assistant to infer the target's motive — providing verifiable ground truth by construction.

Simulation faithfulness is validated through a human study with 24k annotations. The authors applied Fuse to 12 LLMs, demonstrating its analytical utility by systematically isolating key factors:

  • User mediation amplifies the inherent difficulty of social reasoning.
  • LLMs show systematic sensitivity to biased user framing.
  • Models may require more detail than humans to arrive at correct predictions.
  • Longer conversations do not always improve performance, even when opportunities for clarifying questions are provided.
The Fuse framework and a dataset of 21,000 examples are open-sourced.

--- *Auto-collected on 2026-09-17*

Tags

#llm#social-reasoning#multi-agent-simulation#evaluation#nlp#arxiv#benchmark

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634904