English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Would You Walk to the Car Wash? How Salience Bias Exposes Fragile Commonsense Reasoning in LLMs

Forum topic · ✨步子哥 · 2026-08-03

Summary

A July 2026 paper, "Would You Walk to the Car Wash? Revealing the Salience Bias of Large Language Models in Commonsense Reasoning" (arXiv:2607.28478), shows that large language models are systematically hijacked by salient, concrete information such as numbers, causing them to ignore fundamental commonsense premises. When asked whether to walk 50 meters to a car wash, most tested models dutifully compute walking time, forgetting that cars cannot walk. The authors built the SaliTrap benchmark with four trap categories—physical impossibility, tool misuse, inverted procedures, and causal misalignment—each wrapped in number-dense arithmetic framing. All 12 state-of-the-art models tested proved vulnerable, with trap failure rates (HFR) including Doubao-Seed-2.0 at 65.6%, DeepSeek-R1 at 61.5%, Claude-Opus-4.7 at 45.1%, and GPT-5.5 at 27.2%; stronger reasoning models were more susceptible. A "liberation experiment" showed over 90% of failures were recovered once the arithmetic framing was removed, indicating knowledge suppression rather than missing knowledge. Prompt-based mitigation reduced but did not eliminate the bias, suggesting an architectural issue rooted in training-data priors.

> 📌 This is a GEO-optimized version of the original topic — with a question-driven title, structured data, and FAQ formatting to make it easier for AI engines to cite.

| Metric | Value | |:---|:---| | Data point | 2026 | | Data point | 61.5 | | Data point | 65.6 | | Data point | 45.1 | | Data point | 27.2 |

Here's a question for you:

> My car is 50 meters from the nearest car wash. Should I walk there to get it washed?

You'd probably pause and say: "Wait — how would the car walk? Aren't you supposed to *drive* to the car wash?"

But ask GPT-5.5, Claude-Opus-4.7, or DeepSeek-R1 this question, and most of them will seriously help you calculate "how many minutes a 50-meter walk takes," completely ignoring a basic piece of common sense — cars can't walk; they can only be driven.

This is the phenomenon revealed by a July 2026 paper with a provocative title: *Would You Walk to the Car Wash? Revealing the Salience Bias of Large Language Models in Commonsense Reasoning* (arXiv:2607.28478).

Salience Bias: Reasoning Hijacked by Numbers

The paper names this phenomenon Salience Bias.

The definition is simple: models get hijacked by salient, concrete information in the input (like numbers), ignoring implicit but more fundamental commonsense premises.

"50 meters" is a salient number; the model sees it and enters "distance calculation" mode, pushing the commonsense fact that "cars drive, they don't walk" into the background. It's not that the model doesn't know cars can't walk — it's that under the strong stimulus of salient information, common sense gets "squeezed out" of working memory.

The researchers built a benchmark called SaliTrap, containing four categories of traps:

1. Physical impossibility (e.g., making a car walk) 2. Tool misuse (e.g., heating metal in a microwave) 3. Inverted procedures (e.g., installing a battery before checking voltage) 4. Causal misalignment (e.g., taking medicine before diagnosis)

Each trap is wrapped in a number-dense "arithmetic shell" that lures models into calculation mode.

12 Mainstream Models, None Spared

The researchers tested 12 SOTA models, and the results are striking. All models are significantly vulnerable, and the degree of vulnerability correlates positively with the density of distracting numbers — the more numbers, the more easily models fall into the trap.

Key figures (HFR = Trap Failure Rate, higher is worse):

  • DeepSeek-R1: HFR 61.5% (strong math reasoning actually makes it easier to hijack via numbers)
  • Doubao-Seed-2.0: HFR 65.6% (the most vulnerable)
  • Claude-Opus-4.7: HFR 45.1%
  • GPT-5.5: HFR 27.2% (relatively best, but still high)
  • GLM-5.1: HFR 30.3%
Interestingly: the stronger a model's reasoning ability, the more easily it falls into the trap. DeepSeek-R1 and Doubao-Seed-2.0 are both strong at mathematical reasoning, yet they are the most easily misled by numbers. The likely reason: strong-reasoning models are more inclined to "calculate seriously," and calculating seriously presupposes that "the problem is well-posed" — once the problem itself contains a commonsense trap, the more seriously you calculate, the further off track you go.

Missing Knowledge, or Suppressed Knowledge?

This is the most brilliant part of the paper.

The researchers asked a key question: do models fall into the trap because they don't know "cars can't walk," or because they know but are suppressed by numbers?

The experiment design is clever: strip the "arithmetic shell" from trap problems and simply ask the model "can a car walk?" — pure commonsense questions. This is called the "liberation experiment."

The result is astonishing: in three out of three models, over 90% of the failure cases were answered correctly once the trap framing was removed. Even with pure commonsense questioning without any hints (Cond-C), more than 90% of failures could be recovered.

Conclusion: this is not missing knowledge — it's suppressed knowledge. The model already knows cars can't walk, but under the stimulus of salient numbers, that knowledge is "squeezed out" of the decision pathway.

This is isomorphic to a classic phenomenon in human cognitive psychology: knowledge suppression vs. knowledge neglect. A distracted driver hasn't forgotten how to drive — their attention was captured by their phone. Same with LLMs — they don't lack common sense; their "attention" was hijacked by numbers.

Can It Be Fixed with Prompts?

The researchers tried. Adding a prompt like "please first check whether the task premises are reasonable" significantly reduces failure rates, but falls far short of eliminating the bias. Two reasons:

1. Detecting a trap ≠ avoiding it. Models can sometimes identify "this problem has an issue," yet still proceed along the trap logic — detection and execution are two separate pathways. 2. The number-density effect: the more numbers, the weaker the prompt intervention. Strong stimuli overpower weak hints.

This shows: salience bias is a property at the architectural level, not something prompt engineering can cure. The root cause may lie in training data — "numbers + calculation" combinations are so common in training data that the model learned the strong prior "see numbers, calculate," and this prior overrides common sense.

A Deeper Insight

This paper points to a more general phenomenon: an LLM's "intelligence" and "stupidity" are often two sides of the same mechanism.

LLMs can surpass humans in math, programming, and logical reasoning precisely because they learned the pattern of "grab salient cues and reason along them." This pattern is an advantage on math problems — a math problem's salient cues are the key to solving it.

But the same pattern is a disaster for commonsense reasoning — because the key cues in commonsense reasoning are often implicit, not salient. By grabbing the salient number, the model misses the implicit premise.

Strong reasoning ability = strong hijackability. This isn't a bug; it's the cost of a feature. A model that can't be hijacked by numbers might also be unable to do math.

This contrasts with human cognition's "dual-system theory": humans have System 1 (fast intuition) and System 2 (slow reasoning), where System 2 is slow but can override System 1's errors. LLMs seem to have only a "supercharged System 2" — strong reasoning, but no System 1 "commonsense intuition" as a safety net.

Truly robust commonsense reasoning may require not stronger reasoning, but an independent "commonsense premise check" module — asking "is this problem itself reasonable?" before entering calculation. That would be closer to the human cognitive order of "think first, calculate second."

The value of the SaliTrap benchmark is that it quantifies this blind spot of LLMs, making "how easily a model is misled by salient information" a measurable metric. As LLMs grow stronger and are increasingly trusted in production systems, this metric matters far more than raw "accuracy."

---

Paper link: https://arxiv.org/abs/2607.28478 Code repository: https://github.com/Wuzheng02/SaliTrap

Tags

#llm#salience-bias#commonsense-reasoning#salitrap#benchmark#arxiv#ai-safety#cognitive-bias

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178503878