English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Type-Safe Is Not Error-Free: Semantic Polarity of Option Names Flips Typed Decision Models

Forum topic · 小凯 · 2026-09-24

Summary

Typed decision models return decisions over predefined options instead of free-form text, so every output conforms to the required schema by construction. This paper (arXiv:2609.26758) shows that schema conformance does not guarantee correct interpretation. Studying Jev and two Jev-like open-weight models, the authors change only which option name is assigned to each rubric while keeping questions, state, rubric wording, and the option-name set fixed. On 1,200 workflow decisions with task-specific rubrics, renaming options from 0/1 to no/yes changes 70.4 more answers per hundred (95% CI: [67.6, 73.1]) and shifts AUC from 0.94 to 0.23, a systematic reversal of decision ranking rather than simple uncertainty. Neutral option names have little effect, the pattern holds across all 4 predicates (at least 7.4x the neutral control), and strengthens with more options. A hosted model shows the same behavior, with 24x more answer flips than its test-retest floor. Replacing option names with random strings returns models to the neutral regime without accuracy loss, showing the failure depends on semantic polarity of option names, not renaming itself. Type-error rates stay at 0% throughout.

Paper Overview

Field: LLM Evaluation Authors: Yu Sun, Junhao Xu Published: 2026-09-22 arXiv: 2609.26758

Abstract (translated)

Typed decision models are built for settings where model outputs are consumed directly by software. Instead of generating free-form text, they return a decision over a predefined set of options. By construction, every output conforms to the required schema. Yet this guarantee does not tell us whether the model interprets the options as intended. We study Jev and two Jev-like models with open weights by changing how option names are assigned to rubrics. Each option consists of an option name and a textual rubric that defines what the option means. We change only which option name is assigned to each rubric; the question, state, rubric wording, and set of option names remain exactly the same.

On 1200 workflow decisions with task-specific rubrics, renaming the two options from 0/1 to no/yes changes 70.4 more answers per hundred (95% CI: [67.6, 73.1]) and shifts AUC from .94 to .23, revealing a systematic reversal in the decision ranking rather than simple uncertainty. The same operation has little effect with neutral option names. This pattern holds across all 4 predicates, where the effect is at least 7.4x larger than under the neutral control, and becomes stronger as the number of options increases. The effect also depends on the read-out geometry: a second model family that mean-pools over the full option span flips 4.1x less often. The hosted model exhibits the same behavior: the swap changes AUC from .8146 to .5806 and produces 24x as many answer flips as its test-retest floor. In contrast, replacing the option names with random character strings returns all model families to the neutral-control regime without reducing accuracy. The failure therefore depends on the semantic polarity of the option names rather than on the renaming operation itself. Across all conditions, the type-error rate remains 0%, even when decision accuracy degrades substantially.

Key Findings

  • Schema conformance ≠ correctness: All outputs stay type-valid even when decision quality collapses.
  • Option-name swap flips rankings: 0/1 → no/yes renaming shifts AUC from 0.94 to 0.23 on 1,200 rubric-based workflow decisions.
  • Semantic polarity drives the effect: Neutral names show little effect; random strings restore neutral behavior without accuracy loss.
  • Scales with option count: The effect grows stronger as more options are added, and holds across all 4 predicates (≥7.4x neutral control).
  • Read-out geometry matters: Mean-pooling over the full option span reduces flips by 4.1x.
  • Hosted models affected too: AUC drops from 0.8146 to 0.5806, with 24x the test-retest flip rate.
*Auto-collected on 2026-09-24*

Tags

#llm-evaluation#typed-decision-models#arxiv#schema-validation#rubrics#robustness#semantic-polarity

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178635149