Overview
- Field: NLP
- Authors: Haobo Zheng, Tan Tang, Yan Chen, Weijie Wang, Yingcai Wu
- Published: 2026-09-22
- arXiv: 2609.26780
- GroupMemBench / SocialMemBench / EverMemBench binary accuracy: 47.9% / 69.2% / 61.9%
- 62.33% on the publicly reported EverMemBench leaderboard from EverMind-AI — the best reported result among latest state-of-the-art frameworks
- 70.85% on all 1,986 LoCoMo questions (used as a two-person long-term conversation boundary test)
- In a controlled 305-question evaluation, RL raises the SFT Writer's mean accuracy from 57.38% to 68.20%
- Ablations show that the verbatim and structured tracks, as well as person-level and group-level views, are complementary under the standardized evaluation interface
Abstract
Long-term conversational memory in multi-party settings requires more than retrieving relevant content from long-term conversations: it must distinguish who said what, whom each statement concerns, how individuals perceive one another, what information is shared by the group, and how states change over time. Recent studies on multi-party dialogue benchmarks show that existing general-purpose LLM memory systems tend to lose person and group relations or struggle to integrate clues distributed across members, groups, and time. Together, these issues reveal two core bottlenecks: message attribution and relational understanding in multi-party dialogue, and state reconstruction from interleaved histories.To address both, the authors propose SpeakerMem-R1: its dual-track memory stores speaker-labeled verbatim messages and derived states organized into person-level and group-level views, then combines evidence from both tracks by entity, event, and time at query time. To reduce attribution and update errors during structured memory construction while enabling local deployment, Writer-R1 is trained with SpeakerLevenshtein and speaker-conditioned GRPO.
Results
--- *Auto-collected on 2026-09-24*