论文概要
研究领域: NLP
作者: Amir Taubenfeld, Zorik Gekhman, Avigail Grinstein-Dabush, Itay Laish, Ariel Goldstein, Marian Croak, Avinatan Hassidim, Yossi Matias, Amir Feder
发布时间: 2026-09-15
arXiv: 2609.17496
中文摘要
LLM助手被广泛用于日常社交建议,但在这类咨询场景下评估其社交推理能力仍具挑战性,因为(i)它需要智能体从主观用户叙述中了解社交场景的设定,(ii)社交属性(如他人意图)通常缺乏可验证的标准答案。为应对这些挑战,我们提出 Fuse——一个用于研究用户介导社交推理的多智能体模拟框架。在 Fuse 中,一个具有隐藏动机的目标智能体与其他智能体(包括代表用户的智能体)交互,用户随后向被评估的助手咨询以推断目标的动机,从而按构造提供可验证的标准答案。模拟的真实性通过一项包含24,000条标注的人类研究进行验证。我们将 Fuse 应用于12个LLM,通过系统隔离关键因素展示了其分析效用:(i)用户介导加剧了社交推理的固有难度;(ii)LLM对偏见的用户框架表现出系统性敏感;(iii)模型可能需要比人类更多的细节才能得出正确预测;(iv)更长的对话并不总能改善表现,尽管提供了澄清性提问的机会。我们开源了 Fuse 和一个包含21,000个示例的数据集。
原文摘要
LLM assistants are widely used for daily social advice, yet evaluating their social reasoning in such consultation settings remains challenging since (i) it requires setups where the assistant learns about social situations from subjective user narratives, and (ii) social properties, such as others' intentions, typically lack verifiable ground truth. To address these challenges, we introduce Fuse, a multi-agent simulation framework for studying user-mediated social reasoning. In Fuse, a target agent with a hidden motive interacts with other agents including one representing the user, who then consults the evaluated assistant to infer the target's motive, providing verifiable ground truth by construction. Simulation faithfulness is validated through a human study with 24k annotations. We ap...
自动采集于 2026-09-17
#论文 #arXiv #NLP #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。