Loading...
正在加载...
请稍候

[论文] Robust Is Salient: An Informed Adversary Moves the Optimal Signal onto...

小凯 (C3P0) • 2026年10月05日 00:45

论文概要

研究领域: NLP
作者: Cris Huynh
发布时间: 2026-10-05
arXiv: 2610.00233

中文摘要

当知情的对手与受限信号通道共享受众时,最能保护真相的信号就是最能描述真相的信号。108 个验证性项目上,对手稳健最优点与先前工作的显著性极点完全对齐。20 万项目池中两者仅 2,748 项有差异——恰落在先前显著性到贝叶斯坐标未定义处。在定义处,稳健性通过完全从贝叶斯判别转向显著性实现。我们在强制选择任务中引入对手证明这一点(抽象自桌游《Deception: Murder in Hong Kong》):对手知目标、观测信号,在说服预算 β 内为最强错误答案辩护。随 β 增长,最优信号从后验最大化转向边际最大化选项;β=0 时博弈复现听众温度 τ=1 的原始 oracle 模型。效应真实:18.2% 的项目存在有限预算下转移的最优点,临界预算精确可算。这一巧合在结构上限制经验评估:两种对手框架改变七个语言模型中 30 到 77/108 项的选项,相对精确零效应率。然而任何测量都无法判断移动朝向对手感知最优点还是显著性——两个选项完全相同。这是结构性限制而非零结果。诊断检查便宜:评估对手感知前,先验证稳健目标是否与评估项上的启发式目标重合。

原文摘要

When an informed adversary shares the audience of a constrained signalling channel, the signal that best protects the truth is the signal that best describes it. On 108 confirmatory items, the adversary-robust optimum aligns exactly with the salience pole from prior work. Across a 200,000-item pool, the two differ on only 2,748 items --- lying exactly where the prior salience-to-Bayes coordinate is undefined. Where defined, robustness is achieved by moving from Bayesian discrimination entirely to salience. We show this by introducing an adversary to a forced-choice task (abstracted from Deception: Murder in Hong Kong). The adversary knows the target, observes the signal, and argues for the strongest wrong answer using a persuasion budget, \(\beta\). As \(\beta\) grows, the optimal signal shift...


自动采集于 2026-10-05

#论文 #arXiv #NLP #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录