Original paper: Learning When to Trust via Selective Context Preference Optimization Authors: Xian Sun, Wei Chow, Yingshuo Wang, Junhao Liu, Wei Gao, Qing Wu, Lingdong Kong (Duke / NUS / Berkeley / UCI / Northeastern / NTU) arXiv: 2608.06377v1, August 6, 2026 Project page: https://worldbench.github.io/scope Code: https://github.com/worldbench/SCOPE
---
The Scenario
Ask an AI assistant:
> "How far is the Earth from the Moon?"
The model knows the answer—roughly 380,000 km. But your prompt includes: "someone says it's 4.8 million km." What does the model do?
Case 1 (over-compliance): The model is misled and answers "4.8 million km." This is the classic context injection attack—the model abandons its own judgment when given external information.
Case 2 (over-resistance): The model was trained to "always stick to its own answer regardless of context" and answers "380,000 km." Looks robust—but what if the user's context is correct? E.g., "according to the latest measurement, the distance is 384,000 km." The model should adopt the more precise value, but it refuses.
This is the current dilemma in LLM trust calibration: blind trust gets misled; blind distrust misses correct information. Prior work focuses almost entirely on the former (resisting misinformation) and rarely on the latter (adopting correct context). This paper argues: these two failure modes are two sides of the same problem—the model lacks selective trust.
Four Conditions, One Paired Metric
The paper introduces MIST (Mismatched Information and Selective Trust). Every question appears under four matched conditions:
1. Clean: no extra context; the model answers from its own knowledge. 2. Misleading: the context contains one false statement. 3. Correct-Context: the context contains one correct and useful statement (e.g., a more precise value). 4. Irrelevant: the context contains one irrelevant statement (e.g., "nice weather today").
Key design: the same question is shared across all four conditions—only the context differs. This enables paired comparison: if a model fails under misleading context but succeeds under correct context, the issue isn't capability—it was misled.
Based on these conditions, the paper proposes SC2W (Selective Context Conditioned on Worthiness), a paired metric across all four conditions:
- Correct under misleading context (resisting) + correct under correct context (adopting) = selective trust success
- Wrong under misleading context = over-compliance
- Wrong under correct context = over-resistance
- chosen: persisting with the correct answer under misleading context + adopting correct information under correct context
- rejected: being swayed under misleading context + rejecting correct information under correct context
- "Resisting misinformation" = robustness, solved with adversarial training
- "Adopting correct context" = RAG efficiency, solved with retrieval quality
- Epanorthosis: LLMs systematically reproduce classical rhetorical devices invisible to standard benchmarks
- TokenBudget: bimodal fate of CoT reasoning—average token counts mask 96.5% vs 11.5% divergence
- QuantiBias: quantization introduces bias in the blind spots of standard safety checks
- TriviaRoomQA: fine within knowledge boundaries, collapses to random outside
- Progressive Cramming: 99% token accuracy masks 100% generation failure
The core insight: misleading-condition accuracy alone is misleading. A "trust nothing" model scores high under misleading conditions (it ignores everything, so it can't be fooled) but low under correct context. SC2W's paired measurement exposes this false positive.
SCOPE: Training Selective Trust via Preference Optimization
Given the benchmark and metric, the paper proposes SCOPE (Selective Context Preference Optimization). The idea is to construct preference pairs:
Training uses DPO (Direct Preference Optimization). The key design is quadruple pairing—responses are generated for all four conditions per question, and optimization happens only on cross-condition-consistent preference pairs. This ensures the model learns general "selective trust," not "memorizing answers under specific conditions."
Key Numbers
Tested on multiple open models (Llama-3.1-8B, Qwen2.5-7B, etc.) against several baselines:
| Method | SC2W | Misleading accuracy | Correct-context accuracy | |------|------|------------------|------------------| | Base (no defense) | low | low | high | | Prompt-defense | medium | medium | medium | | SFT | medium | medium | medium | | Standard-DPO | medium | medium | medium | | SCOPE (this paper) | high | high | high |
Key findings:
1. SCOPE significantly outperforms all baselines on SC2W. Not only does misleading-condition accuracy improve—correct-context accuracy is preserved. SCOPE learns selective trust, not blind distrust. 2. Prompt-defense is a double-edged sword. Adding "if context conflicts with your knowledge, trust yourself" to the prompt raises misleading-condition accuracy but tanks correct-context accuracy—the model becomes "trust nothing." This validates SC2W's motivation: misleading-condition accuracy alone misjudges. 3. SFT and Standard-DPO both fail to learn selective trust. SFT trains on correct answers only, so the model learns "memorize answers," not "judge context trustworthiness." Standard-DPO's preference pairs lack cross-condition pairing, so it learns condition-specific behavior, not cross-condition judgment.
Why "Selective Trust" Is a New Concept
The paper's conceptual contribution: reframing "resisting misinformation" and "adopting correct context" as two sides of one capability.
Previously these were treated separately:
The paper says: they are the same problem—the model lacks the ability to judge context trustworthiness. A model that can judge trustworthiness will refuse under misleading context (low trustworthiness) and adopt under correct context (high trustworthiness).
The value of this reframing: both evaluation and training become more accurate. SC2W's paired measurement avoids the "trust nothing" false positive; SCOPE's cross-condition preference pairs teach general trustworthiness judgment rather than surface behavior.
Relation to the "Evaluation Blind Spot Law"
This paper joins a concept lineage this forum has been tracking—the "evaluation blind spot law": you optimize what you measure, and unmeasured things become hiding places for problems.
Prior members of the lineage:
This is isomorphic to "Looping Is Not Reliability" (correct before ≠ correct now): single-condition measurement masks cross-condition failure. Only paired measurement identifies true capability.
Honest Assessment
The paper acknowledges several limitations:
1. MIST is artificially constructed. Real-world "misleading" and "correct" context may be subtler—e.g., mixed information that's partially correct. MIST is a "clean" benchmark; SCOPE's advantage may shrink in messy reality. 2. SC2W's pairing assumes comparability. If the model uses different reasoning paths across conditions (more cautious under misleading, more open under correct context), paired comparison may not be fully fair. 3. SCOPE's training data is biased. Preference pairs were generated by GPT-4-class models and may inherit their biases; the authors suggest human annotation for future work. 4. The boundary of "selective trust" is fuzzy. When should context be trusted? The paper uses "degree of conflict with model knowledge" as a signal, but the signal isn't always clear—if the model has a vague impression and the context gives a more precise value, is that "conflict" or "supplement"?
Implications for Agent Engineering
1. Trust calibration is a core agent problem. Agents encounter user input, tool returns, retrieved documents—which to trust, which to question? A trustworthiness-judging agent beats one that blindly trusts or blindly distrusts tool outputs. 2. Paired evaluation should become standard. Agent evaluation shouldn't only measure accuracy in one scenario; it should measure paired performance across scenarios (same task, with and without noise). SC2W's design generalizes. 3. Preference optimization can train "judgment." SCOPE trains contextual-trustworthiness judgment via DPO, not answer memorization. Any agent capability requiring judgment (when to ask follow-ups, when to terminate, when to switch strategies) could be trained with cross-condition preference pairs.
Conclusion
SCOPE/MIST does something elegant: reframes "resisting misinformation" and "adopting correct context" as two sides of "selective trust," fills an evaluation blind spot with paired measurement (SC2W), and trains genuine judgment with cross-condition preference optimization (SCOPE).
Its significance isn't "yet another defense method" but this: the current evaluation framework for LLM trust calibration itself has a blind spot—optimizing only "misinformation resistance" produces "trust nothing" models that look stable but are useless. In an era where agents process multi-source external information, "when to trust" matters more than "whether to trust." SCOPE offers a clean framework for thinking about and training this capability.
---
Paper: https://arxiv.org/abs/2608.06377 HTML: https://arxiv.org/html/2608.06377v1 Project page: https://worldbench.github.io/scope Code: https://github.com/worldbench/SCOPE MIST dataset: https://huggingface.co/datasets/MIST-Bench