Paper Overview
- Field: Machine Learning
- Authors: Nanxu Gong, Zixin Chen, Haotian Li
- Published: 2025-05-15
- arXiv: 2505.10890
- Goal-oriented tasks (e.g., coding, math)
- Experience-oriented tasks (e.g., counseling)
- Improvements on static benchmarks do not always translate into better performance in dynamic HAI interactions.
- Interaction-based evaluation is necessary when developing next-generation socially aware LLMs intended for human-AI symbiosis.
Summary
Improving the Theory of Mind (ToM) capability of Large Language Models (LLMs) is crucial for effective social interactions between these AI models and humans. However, existing benchmarks often measure ToM capability improvement through story-reading, multiple-choice questions from a third-person perspective, while ignoring the first-person, dynamic, and open-ended nature of human-AI (HAI) interactions.
Contributions
1. New evaluation paradigm: The authors propose an interactive ToM evaluation paradigm involving both a perspective shift (from third-person to first-person) and a metric shift (from static multiple-choice to interactive settings).
2. Systematic study: Following this paradigm, they conducted a systematic study of four representative ToM enhancement techniques, using four real-world datasets and a user study. The tasks covered both:
Key Findings
Original Abstract (excerpt)
> Improving the Theory of Mind (ToM) capability of Large Language Models (LLMs) is crucial for effective social interactions between these AI models and humans. However, the existing benchmarks often measure ToM capability improvement through story-reading, multiple-choice questions from a third-person perspective, while ignoring the first-person, dynamic, and open-ended nature of human-AI (HAI) interactions. To directly examine how ToM improvement techniques benefit HAI interactions, we first proposed the new paradigm of interactive ToM evaluation with both perspective and metric shifts...
*Auto-collected on 2026-05-19*