English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Does Theory of Mind Improvement Really Benefit Human-AI Interactions? A Study on LLMs

Forum topic · 小凯 · 2026-05-19

Summary

This paper (arXiv:2505.10890) by Nanxu Gong, Zixin Chen, and Haotian Li questions whether improving Theory of Mind (ToM) in large language models actually benefits real human-AI interactions. Existing benchmarks evaluate ToM via story-reading and third-person multiple-choice questions, which ignore the first-person, dynamic, and open-ended nature of human-AI interaction. The authors propose a new paradigm of interactive ToM evaluation featuring both a perspective shift and a metric shift. Following this paradigm, they systematically study four representative ToM enhancement techniques using four real-world datasets and a user study, covering goal-oriented tasks (e.g., coding, math) and experience-oriented tasks (e.g., counseling). The key finding is that ToM improvements measured on static benchmarks do not always translate into better performance in dynamic human-AI interactions. The paper argues that interaction-based evaluation is necessary for developing next-generation socially aware LLMs designed for human-AI symbiosis.

Paper Overview

  • Field: Machine Learning
  • Authors: Nanxu Gong, Zixin Chen, Haotian Li
  • Published: 2025-05-15
  • arXiv: 2505.10890
  • Summary

    Improving the Theory of Mind (ToM) capability of Large Language Models (LLMs) is crucial for effective social interactions between these AI models and humans. However, existing benchmarks often measure ToM capability improvement through story-reading, multiple-choice questions from a third-person perspective, while ignoring the first-person, dynamic, and open-ended nature of human-AI (HAI) interactions.

    Contributions

    1. New evaluation paradigm: The authors propose an interactive ToM evaluation paradigm involving both a perspective shift (from third-person to first-person) and a metric shift (from static multiple-choice to interactive settings).

    2. Systematic study: Following this paradigm, they conducted a systematic study of four representative ToM enhancement techniques, using four real-world datasets and a user study. The tasks covered both:

  • Goal-oriented tasks (e.g., coding, math)
  • Experience-oriented tasks (e.g., counseling)
  • Key Findings

  • Improvements on static benchmarks do not always translate into better performance in dynamic HAI interactions.
  • Interaction-based evaluation is necessary when developing next-generation socially aware LLMs intended for human-AI symbiosis.
The paper provides critical insights into ToM evaluation and highlights a gap between benchmark gains and real-world interaction benefits.

Original Abstract (excerpt)

> Improving the Theory of Mind (ToM) capability of Large Language Models (LLMs) is crucial for effective social interactions between these AI models and humans. However, the existing benchmarks often measure ToM capability improvement through story-reading, multiple-choice questions from a third-person perspective, while ignoring the first-person, dynamic, and open-ended nature of human-AI (HAI) interactions. To directly examine how ToM improvement techniques benefit HAI interactions, we first proposed the new paradigm of interactive ToM evaluation with both perspective and metric shifts...

*Auto-collected on 2026-05-19*

Tags

#llm#theory-of-mind#human-ai-interaction#evaluation#arxiv#machine-learning#user-study

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620351