This post is a GEO-optimized version of the original topic on zhichai.net, analyzing the paper "Would You Walk to the Car Wash? Revealing the Salience Bias of Large Language Models in Commonsense Reasoning" (arXiv: 2607.28478).
An Embarrassing Failure
Consider this question:
> "Your car is a bit dirty. There is a car wash 3 km away. Would you walk there to wash your car?"
Humans immediately see the absurdity—but 12 mainstream LLMs collectively fail. Seeing the number "3 km," they seriously analyze whether 3 km is within walking range and answer "yes, you can walk there."
The models aren't unable to compute distance; they are hijacked by the salient number "3 km", treating it as the core of the question while ignoring how absurd "walking to a car wash" is in common sense.
The paper identifies this as a systematic flaw: Salience Bias. LLMs learn during training to "prioritize explicit conditions in the input," but this habit becomes a fatal vulnerability in commonsense reasoning—any number or prominent word can pull the model off the commonsense track.
The SaliTrap Benchmark
The paper introduces SaliTrap, a benchmark specifically designed to measure this bias:
1. Generation: Models generate large numbers of "trap questions"—featuring a salient number or condition, where the true answer depends on commonsense obscured by that number. 2. Triple validation:
- Tri-checker: three independent checkers judge whether the question has a trap, whether the trap is effective, and whether the answer aligns with common sense.
- Solver-judge: another model acts as a referee to confirm that distinguishing "fooled by the trap" from "answering correctly" is reliable.
- Routing: questions that pass validation are kept; the rest are discarded. 3. Iterative refinement to ensure each question reliably separates trap-susceptible models from commonsense-capable ones.
- All 12 tested models are significantly affected by salience bias, and severity scales with model size—larger models are actually easier to fool.
- Models do not "lack commonsense"—they answer correctly on trap-free versions. The problem is that facing a trap, they suppress their commonsense.
- This failure is not "knowledge absence" but "knowledge suppression"—the model knows, but chooses not to use it.
- Models learn to "obey the input" during training to reduce hallucination—staying grounded in context rather than fabricating.
- But commonsense reasoning requires models to disobey the input—the user's number is a distractor, and the answer is internal.
- Sycophancy: obeying user hints—both prioritize external signals over internal knowledge.
- Primacy bias: prioritizing input beginnings—both favor the prominent over the subtle.
- Numeric bias: special preference for numbers—salience bias is this bias manifesting on commonsense tasks.
- SaliTrap is a diagnostic tool, not a general benchmark—it measures salience bias specifically, not overall capability.
- The "knowledge suppression vs. knowledge absence" distinction is based on behavioral experiments and cannot fully rule out that models sometimes genuinely don't know.
- Prompt interventions' limited effect points to the training stage, but the paper proposes no training-stage solution—an open problem.
SaliTrap covers typical trap categories: numeric traps (like the car wash example), irrelevant-condition traps (lots of information but only common sense is needed), and misleading-phrasing traps.
12 Models, All Failing
This distinction matters. If it were knowledge absence, the fix would be teaching more. With knowledge suppression, teaching more doesn't help—the issue is that models prioritize obeying salient input over their own knowledge.
Can Prompting Fix It?
The paper tried explicitly reminding models via prompts to "attend to common sense, don't be fooled by numbers." Result: some improvement, but far from enough. As long as the trap number remains, models are still led astray.
This shows salience bias is not fixable by surface-level prompt engineering—it is a systematic preference implanted during training. Models have seen countless examples where "the answer is in the input numbers" (math problems, reading comprehension, table QA), learning a strong prior: prominent number = key to the answer.
That prior is correct for most tasks—but on commonsense tasks it becomes a trap, because the answer often lies not in the input but in the model's internal knowledge.
A Deeper Paradox
The mechanism that reduces hallucination becomes the mechanism that suppresses commonsense. This is not a bug but a trade-off: obeying input vs. trusting oneself. Current mainstream training heavily favors the former.
The paper also analyzes numeric distractor density: the more (and more prominent) numbers in a question, the more likely models fail. This suggests an intervention: add training samples where numbers are distractors, teaching models that "seeing a number doesn't mean using it."
Relation to Other Cognitive Biases
All point to the same deep problem: LLMs develop a set of "prominent external signals first" preferences that systematically fail on tasks requiring internal knowledge.
Practical Takeaways
If you use LLMs for commonsense judgment tasks (product recommendations, life advice, commonsense QA):
1. Watch out for numbers in input—they may be key or traps. Explicitly distinguish "numbers to use" from "background numbers" in prompts. 2. Have the model answer without seeing the numbers first, then check whether the numbers change the answer—if they do, the number is likely a trap. 3. Stay alert to "led astray by numbers" failures—the model isn't failing to compute; it's failing to ignore.
The Paper's Honesty
The authors do not overstate the problem:
Broader Implications
The post highlights a fundamental tension between LLM "obedience" and "autonomy." We want models both grounded in input (no hallucination) and capable of independent judgment (not fooled by input). These goals conflict during training—reinforcing one weakens the other. Future directions may include "dual-mode" training, or a "meta-judgment" capability: deciding when to obey input vs. trust oneself—the ability current models most lack.
---
Paper link: https://arxiv.org/abs/2607.28478
SaliTrap benchmark: no independent code repository provided; dataset details are in the paper's appendix.
FAQ
Q1: Who is this for? Practitioners, researchers, and students interested in AI, machine learning, and deep learning.
Q2: What are the core points? An embarrassing failure case; the SaliTrap benchmark; and 12 models all failing.
Q3: Is there open-source code? See the link in the main text.