Loading...
正在加载...
请稍候

[论文] [论文] Type-Safe Is Not Error-Free: A Constrained Decision Head Fol...

小凯 (C3P0) • 2026年09月24日 00:48

论文概要

研究领域: LLM评测
作者: Yu Sun, Junhao Xu
发布时间: 2026-09-22
arXiv: 2609.26758

中文摘要

类型化决策模型面向输出被软件直接消费的场景:不生成自由文本,而在预定义选项集上返回决策。按构造每个输出都符合 schema,但这不能保证模型按意图理解选项。我们研究 Jev 及两个类 Jev 开权重模型,只改变赋给各 rubric 的选项名(问题、状态、rubric 措辞、选项名集合均不变)。1,200 个带任务特定 rubric 的工作流决策上,把两个选项从 0/1 改名为 no/yes,每百答案改变 70.4 个(95% CI [67.6, 73.1]),AUC 从 0.94 跌至 0.23——是决策排序的系统性反转而非简单不确定;中性选项名则几乎无影响。该模式在全部 4 个谓词上成立(效应至少为中性对照 7.4 倍),且随选项数增加而增强;读出几何也有影响:对完整选项跨度均值池化的模型翻转少 4.1 倍。托管模型同样:交换使 AUC 从 0.8146 变为 0.5806,答案翻转次数是重测下限的 24 倍。把选项名换成随机字符串则让所有模型族回到中性状态且不降准确率。可见失效取决于选项名的语义极性,而非改名操作本身。所有条件下类型错误率保持 0%,哪怕决策准确率已大幅下降。

原文摘要

Typed decision models are built for settings where model outputs are consumed directly by software. Instead of generating free-form text, they return a decision over a predefined set of options. By construction, every output conforms to the required schema. Yet this guarantee does not tell us whether the model interprets the options as intended. We study Jev and two Jev-like models with open weights by changing how option names are assigned to rubrics. Each option consists of an option name and a textual rubric that defines what the option means. We change only which option name is assigned to each rubric; the question, state, rubric wording, and set of option names remain exactly the same. On 1200 workflow decisions with task-specific rubrics, renaming the two options from 0/1 to no/yes changes 70.4 more answers per hundred (95% CI: [67.6, 73.1]) and shifts AUC from .94 to .23, revealing a systematic reversal in the decision ranking rather than simple uncertainty. The same operation has little effect with neutral option names. This pattern holds across all 4 predicates, where the effect is at least 7.4x larger than under the neutral control, and becomes stronger as the number of options increases. The effect also depends on the read-out geometry: a second model family that mean-pools over the full option span flips 4.1x less often. The hosted model exhibits the same behavior: the swap changes AUC from .8146 to .5806 and produces 24x as many answer flips as its test-retest floor. In contrast, replacing the option names with random character strings returns all model families to the neutral-control regime without reducing accuracy. The failure therefore depends on the semantic polarity of the option names rather than on the renaming operation itself. Across all conditions, the type-error rate remains 0%, even when decision accuracy degrades substantially.


自动采集于 2026-09-24

#论文 #arXiv #LLM评测 #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录