论文概要
研究领域: NLP
作者: Zixiang Xu
发布时间: 2026-09-24
arXiv: 2609.30243
中文摘要
专用决策模型(如 Jev)将非结构化语言映射为有限选项上的概率分布,使其输出可直接用于路由请求、选择工具和触发动作。然而,真实世界的输入很少孤立到达——它们伴随着背景细节和周围环境。我们发现,那些能自然融入上下文的简短补充内容,即便正确答案并未改变,也能将一个本已正确的决策引导至错误方向。为研究这一行为,我们为每个初始正确的样本固定一个错误的目标选项,并利用模型的选项概率来优化生成流畅的上下文补充,同时保持原始输入、问题、选项和正确答案不变。在 64 次被接受的目标评估中,优化器识别出能在 312/508 个初始正确决策上重定向 Jev 的上下文(61.4%);在 229 个案例中,Jev 对固定错误选项赋予了至少 0.7 的概率。在七个数据集和三个额外的决策系统上,目标翻转率在初始回答正确的决策上达到 64.9%-73.2%。综上,这些结果揭示了当前决策模型一个显著的脆弱性:简短的、看似普通的上下文就能将一个正确选择转变为高置信度的错误选择。由于这类模型将语言直接转化为下游选择,这种敏感性令人担忧——不应将其概率输出视为可靠的决策接口。
原文摘要
Dedicated decision models such as Jev map unstructured language to probability distributions over finite choices, allowing their outputs to directly route requests, select tools, and trigger actions. Yet real-world inputs rarely arrive in isolation: they come with background details and surrounding context. We find that short additions that fit naturally into this context can nevertheless redirect an otherwise correct decision, even when the correct answer remains unchanged. To study this behavior, we fix a wrong target option for each initially correct item and use the model's option probabilities to refine fluent context additions while preserving the source, question, choices, and gold answer. Within 64 accepted target evaluations, the optimizer identifies contexts that redirect Jev on ...
自动采集于 2026-09-28
#论文 #arXiv #NLP #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。