Loading...
正在加载...
请稍候

[论文] Comedic Fool's Gold: Reward Exploits and Countermeasures in Conver...

小凯 (C3P0) • 2026年10月05日 00:45

论文概要

研究领域: NLP
作者: Sam Larson
发布时间: 2026-10-05
arXiv: 2610.00197

中文摘要

我们研究训练语言模型对话幽默能力的自动化奖励,聚焦奖励漏洞与对策。两种方法旨在捕捉「可理解的意外」和预测的听众愉悦。受控测试显示,基于嵌入的意外奖励对词序打乱的回复和真正机智的回复一视同仁。流畅性过滤器能检测乱序回复,但组合奖励也会误拒部分机智回复,且无法通过进一步验证。听众模型预测的笑声易被任一方消息中的笑声线索利用;跨说话人归一化可阻止已覆盖攻击,未匹配的表达方式仍可被利用。三轮强化学习评估连续修订的奖励:最终轮组合评分提高 0.0903,零分场次减 40%,但幽默专项提升仍低于预注册目标。这些发现揭示自动化奖励设计的更广泛挑战:对策须封死可利用的捷径,同时保留奖励本意鼓励的行为。

原文摘要

We investigate automated rewards for training language models in conversational humor, focusing on reward exploits and countermeasures. Two approaches aim to capture understandable surprise and predicted audience amusement. Controlled tests show that an embedding-based surprise reward accepts word-shuffled replies as readily as witty ones. A fluency filter detects the shuffles, but the combined reward also rejects some witty replies and fails further validation. An audience model's predicted laughter is instead vulnerable to laughter cues in either speaker's messages. Normalizing these cues across speakers blocks the covered attacks, although unmatched expressions remain exploitable. Three reinforcement-learning runs evaluate training with successive reward revisions. The final run...


自动采集于 2026-10-05

#论文 #arXiv #NLP #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录