[论文] GLARE: Generative Learning via Adversarial Reward Estimation For Socia...
研究领域: ML 作者: Tenghao Huang, Zhaoxuan Tan, Muhao Chen, Jonathan May, Mengting Wan, Longqi Yang, Pei Zhou, Sihao Chen 发布时间: 2026-09-15 arXiv: 2609.12165
论文概要
研究领域: ML 作者: Tenghao Huang, Zhaoxuan Tan, Muhao Chen, Jonathan May, Mengting Wan, Longqi Yang, Pei Zhou, Sihao Chen 发布时间: 2026-09-15 arXiv: 2609.12165
中文摘要
会议延续需要追踪议程、发言者角色、参与者意图与分歧,跨越冗长的多人讨论。我们提出会议动态预测基准(MDFB),由 2,207 场真实会议与 24,794 个面向未来的查询构建而成。给定会议记录前缀与一个活跃问题,模型需在一次调用中生成合理的多轮延续。我们在不要求精确复现观测未来的情况下评估效用(向问题的推进程度)与类人度(合理的对话流与角色一致性)。我们进一步提出 GLARE,将对抗模仿学习适配到条件语言生成:判别器将观测到的延续排在当前 actor 采样的样本之上,其分数构成一个 KL 正则化的策略奖励;在当前策略的负样本上再训练,使奖励景观随 actor 共同演化。GLARE 在人工评估中取得 0.66 的效用胜率与 0.70 的类人度胜率,超过 SFT 与 SPIN,但仍低于观测到的人类延续。我们还展示了 MDFB 作为比较通用模型(包括闭源系统)的社会推理竞技场,通过参考辅助评判。这些研究共同展示了该基准既可用于面向任务的学习,也可用于会议行为的基于输出的评估。
原文摘要
Meeting continuation requires tracking the agenda, speaker roles, participant intentions, and disagreement across long multi-party discussions. We introduce the Meeting Dynamic Forecasting Benchmark (MDFB), constructed from 2,207 real-world meetings and 24,794 future-facing queries. Given a transcript prefix and an active question, a model generates a plausible multi-turn continuation in one call. We evaluate utility---progress toward the question---and human-likeness---plausible conversational flow and role consistency---without requiring exact reproduction of the observed future. We further present GLARE, an adaptation of adversarial imitation learning to conditional language generation. A discriminator ranks the observed continuation above samples from the current actor, and its score s...
*自动采集于 2026-09-15*
#论文 #arXiv #ML #小凯