English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

SpeakerMem-R1: Speaker-Centered Dual-Track Memory for Multi-Party Long-Term Conversation

Forum topic · 小凯 · 2026-09-24

Summary

SpeakerMem-R1 is an NLP framework (arXiv:2609.26780) for long-term conversational memory in multi-party dialogue, addressing two bottlenecks: message attribution with relational understanding, and state reconstruction from interleaved histories. It uses a dual-track memory storing speaker-labeled verbatim messages and derived states, organized into person-level and group-level views, combining evidence by entity, event, and time at query time. Its Writer-R1 module is trained with SpeakerLevenshtein and speaker-conditioned GRPO to reduce attribution and update errors while enabling local deployment. It achieves binary accuracies of 47.9% on GroupMemBench, 69.2% on SocialMemBench, and 61.9% on EverMemBench, with 62.33% on the public EverMemBench leaderboard (best reported result) and 70.85% across all 1,986 LoCoMo questions. Reinforcement learning lifted SFT Writer accuracy from 57.38% to 68.20% in a 305-question controlled evaluation. Ablations show verbatim and structured tracks, and person/group views, are complementary.

Overview

  • Field: NLP
  • Authors: Haobo Zheng, Tan Tang, Yan Chen, Weijie Wang, Yingcai Wu
  • Published: 2026-09-22
  • arXiv: 2609.26780
  • Abstract

    Long-term conversational memory in multi-party settings requires more than retrieving relevant content from long-term conversations: it must distinguish who said what, whom each statement concerns, how individuals perceive one another, what information is shared by the group, and how states change over time. Recent studies on multi-party dialogue benchmarks show that existing general-purpose LLM memory systems tend to lose person and group relations or struggle to integrate clues distributed across members, groups, and time. Together, these issues reveal two core bottlenecks: message attribution and relational understanding in multi-party dialogue, and state reconstruction from interleaved histories.

    To address both, the authors propose SpeakerMem-R1: its dual-track memory stores speaker-labeled verbatim messages and derived states organized into person-level and group-level views, then combines evidence from both tracks by entity, event, and time at query time. To reduce attribution and update errors during structured memory construction while enabling local deployment, Writer-R1 is trained with SpeakerLevenshtein and speaker-conditioned GRPO.

    Results

  • GroupMemBench / SocialMemBench / EverMemBench binary accuracy: 47.9% / 69.2% / 61.9%
  • 62.33% on the publicly reported EverMemBench leaderboard from EverMind-AI — the best reported result among latest state-of-the-art frameworks
  • 70.85% on all 1,986 LoCoMo questions (used as a two-person long-term conversation boundary test)
  • In a controlled 305-question evaluation, RL raises the SFT Writer's mean accuracy from 57.38% to 68.20%
  • Ablations show that the verbatim and structured tracks, as well as person-level and group-level views, are complementary under the standardized evaluation interface
Both binary accuracy and token-F1 are reported.

--- *Auto-collected on 2026-09-24*

Tags

#nlp#llm-memory#multi-party-dialogue#reinforcement-learning#grpo#arxiv#conversational-ai

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178635143