Overview
- Field: NLP
- arXiv: 2609.15972
- Posted: 2026-09-14 (auto-collected 2026-09-16)
- Authors: Zixuan Wang, Yufan Zhou, Jinzhou Tang, Xinle Yu, Chengjun Wu, Lyumanshan Ye, Zhaoxiang Feng, Letian Peng, Adyasha Patra, Fan Bai, Enze Ma, Zhengding Hu, Jianyang Gu, Zhao Wang, Yufei Ding, Jingbo Shang, Tianmin Shu, Zhiting Hu, Zhen Wang
- Psychology-guided simulator: preserves personal characteristics while updating mental states through interaction, generating coherent dialogue.
- Shared evolving mental state: drives both user behavior and guides an Oracle assistant's informed responses.
- Privileged distillation: trains the model to learn the Oracle's informed responses, without requiring direct access to user mental states at deployment.
- Evaluation combines personalization and theory-of-mind to assess human-aware learning.
- Trained on the full Mind2Dialogue corpus, the model outperforms instruction-tuned Qwen, Llama, and OLMo baselines on all reported personalization metrics.
- Preference-following generation improves by 26.6 to 40.9 percentage points.
- Gains extend to belief and action reasoning for Qwen and Llama, beyond personalized assistant behavior.
Abstract (translated)
As language models become more capable, long-term collaboration in learning, reasoning, and decision-making calls for a deeper understanding of the people they serve. Yet training such human-aware language models faces a fundamental supervision gap: current datasets for LLM assistant training contain few if any well-informed responses explicitly grounded in users' unspoken beliefs and goals. Scaling such supervision is inherently constrained, as users' underlying states are not directly observable.
The authors propose the Mind2Dialogue framework to mitigate this gap by simulating users' mental states and turning them into privileged supervision for human-aware training. Key components:
Key results
Outlook
Mind2Dialogue aims to make user simulation a foundation for true AI collaborators that understand the beliefs and intentions behind people's words and support their long-term goals in education, work, and daily life.
Original abstract (excerpt)
> As language models become more capable, long-term collaboration in learning, reasoning, and decision-making calls for a deeper understanding of the people they serve. Yet training such human-aware language models faces a fundamental supervision gap because current datasets for LLM assistant training contain few if any well-informed responses explicitly grounded in users' unspoken beliefs and goals...
---
*Auto-collected on 2026-09-16. Source: arXiv:2609.15972*