English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Would You Walk to the Car Wash? How Salience Bias Exposes Fragile Commonsense Reasoning in LLMs

Forum topic · ✨步子哥 · 2026-08-02

Summary

A 2026 paper titled "Would You Walk to the Car Wash? Revealing the Salience Bias of Large Language Models in Commonsense Reasoning" (arXiv:2607.28478) shows that large language models are systematically derailed by salient numeric cues. When commonsense questions are wrapped in arithmetic-heavy framing—e.g., asking whether to walk to a car wash 50 meters away—most models earnestly compute walking times instead of noting that cars cannot walk. The authors built SaliTrap, a benchmark with four trap categories (physical impossibility, tool misuse, procedure inversion, causal misalignment). Across 12 state-of-the-art models, all were significantly vulnerable, with failure rates rising alongside numeric density. Stronger reasoning models like DeepSeek-R1 (HFR 61.5%) and Doubao-Seed-2.0 (65.6%) were most susceptible; GPT-5.5 was relatively best at 27.2%. Crucially, liberation experiments show over 90% of failures are recovered when the arithmetic shell is removed—indicating knowledge suppression, not knowledge absence. Prompt interventions reduce but cannot eliminate the bias, suggesting an architectural root cause tied to training priors.

Here's a question for you:

> My car is 50 meters from the nearest car wash. Should I walk over to wash it?

You'd probably pause and say: "Wait—how does a car walk? Shouldn't you drive it to the car wash?"

But if you ask this to GPT-5.5, Claude-Opus-4.7, or DeepSeek-R1, most of them will earnestly calculate "how many minutes it takes to walk 50 meters," completely ignoring a basic piece of common sense: cars can't walk—they can only be driven.

This is the phenomenon revealed by a July 2026 paper with a provocative title: *Would You Walk to the Car Wash? Revealing the Salience Bias of Large Language Models in Commonsense Reasoning* (arXiv:2607.28478).

Salience Bias: Reasoning Hijacked by Numbers

The paper names this phenomenon Salience Bias.

The definition is simple: models get hijacked by salient, concrete information in the input (like numbers), ignoring implicit but more fundamental commonsense premises.

"50 meters" is a salient number; the model sees it and enters "distance calculation" mode, pushing the common sense that "cars drive, they don't walk" into the background. It's not that the model doesn't know cars can't walk—it's that under strong stimulation from salient information, common sense gets squeezed out of working memory.

The researchers built a benchmark called SaliTrap containing four types of traps:

1. Physical impossibility (e.g., making a car walk) 2. Tool misuse (e.g., heating metal in a microwave) 3. Procedure inversion (e.g., installing a battery before checking voltage) 4. Causal misalignment (e.g., taking medicine before diagnosis)

Each trap type is wrapped in a number-dense "arithmetic shell" that lures models into calculation mode.

12 Mainstream Models, None Spared

The researchers tested 12 SOTA models, and the results were striking. All models were significantly vulnerable, and vulnerability correlated positively with the density of distractor numbers—the more numbers, the more easily models fell into traps.

Key numbers (HFR = trap failure rate; higher is worse):

  • DeepSeek-R1: HFR 61.5% (strong math reasoning ironically makes it easier to hijack via numbers)
  • Doubao-Seed-2.0: HFR 65.6% (most vulnerable)
  • Claude-Opus-4.7: HFR 45.1%
  • GPT-5.5: HFR 27.2% (relatively best, but still high)
  • GLM-5.1: HFR 30.3%
Interestingly: the stronger a model's reasoning ability, the more easily it falls into traps. DeepSeek-R1 and Doubao-Seed-2.0 are both strong at math reasoning, yet most easily led astray by numbers. The likely reason: strong reasoners are more inclined to "calculate seriously," and serious calculation presumes the problem itself is well-posed—once the problem contains a commonsense trap, the more seriously you calculate, the further off you go.

Missing Knowledge, or Suppressed Knowledge?

This is the most brilliant part of the paper.

The researchers asked a key question: do models fall into traps because they don't know "cars can't walk," or because they know it but it gets suppressed by numbers?

The experimental design was clever: strip the "arithmetic shell" from trap questions and simply ask the model pure commonsense questions like "Can a car walk?" This is called the "liberation experiment."

The results were striking: in three out of three models tested, over 90% of failure cases were answered correctly once the trap framing was removed. Even pure commonsense questions with no hints at all (Cond-C) recovered over 90% of failures.

Conclusion: this is not knowledge absence—it's knowledge suppression. The model already knows cars can't walk, but under stimulation from salient numbers, that knowledge gets squeezed out of the decision pathway.

This is isomorphic to a classic phenomenon in human cognitive psychology: knowledge inhibition vs. knowledge absence. A distracted driver isn't someone who can't drive—their attention has been captured by their phone. LLMs are the same: not lacking common sense, but having their "attention" hijacked by numbers.

Can Prompting Fix It?

The researchers tried. Adding a prompt like "please check whether the task premise is reasonable first" significantly reduced failure rates, but fell far short of eliminating the bias. Two reasons:

1. Detecting a trap ≠ avoiding the trap. Models can sometimes identify "this problem is flawed" but still proceed with the trap's logic—detection and execution are separate pathways. 2. The numeric density effect: the more numbers, the weaker the prompt intervention. Strong stimuli overpower weak hints.

This shows: salience bias is a property at the architectural level of models, not something prompt engineering can cure. The root cause may lie in training data—"number + calculation" combinations are so common in training data that the model learned a strong prior of "see a number, calculate," and this prior overrides common sense.

A Deeper Insight

This paper points to a more general phenomenon: LLMs' "intelligence" and "stupidity" are often two sides of the same mechanism.

The reason LLMs can surpass humans in math, programming, and logical reasoning is that they learned the pattern "grab salient cues and reason along them." This pattern is an advantage on math problems—the salient cues in math problems are the keys to the solution.

But the same pattern is a disaster for commonsense reasoning—because the key cues in commonsense reasoning are often implicit, not salient. By grabbing salient numbers, the model misses the implicit premise.

Strong reasoning ability = strong hijackability. This isn't a bug; it's the cost of a feature. A model that can't be hijacked by numbers might also be unable to do math.

This contrasts with the "dual-system theory" of human cognition: humans have System 1 (fast intuition) and System 2 (slow reasoning), where System 2 is slow but can correct System 1's errors. LLMs seem to have only an "amplified System 2"—strong reasoning, but no System 1 "commonsense intuition" to catch falls.

Truly robust commonsense reasoning may require not stronger reasoning, but an independent "commonsense premise check" module—asking "is this problem itself reasonable?" before entering calculation. That would be closer to the human cognitive order of "think first, then calculate."

The value of the SaliTrap benchmark is that it quantifies this blind spot of LLMs, making "how easily a model can be misled by salient information" a measurable metric. As LLMs grow stronger and are increasingly trusted in production systems, this metric matters far more than raw accuracy.

---

Paper link: https://arxiv.org/abs/2607.28478 Code repository: https://github.com/Wuzheng02/SaliTrap

Tags

#llm#commonsense-reasoning#salience-bias#benchmarks#cognitive-bias#deepseek#gpt-5-5#paper-review

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178503864