English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Language Models Compare Quantities Using Number-Specific and Unit-Specific Heuristics Rather Than Shared-Scale Conversion

Forum topic · 小凯 · 2026-06-04

Summary

This arXiv paper (2606.03982) investigates how language models (LMs) compare quantities with measurement units, such as 110 cm versus 1.2 m, which requires combining a numeral with a symbolic unit scale. Studying controlled comparison tasks across multiple unit systems, the authors find that accuracy degrades near the comparison boundary, where small changes in value determine the correct answer. The errors are systematic: linear surrogate models predict LM preferences from numerical-difference and unit-scale-difference cues, and causal interventions on subspaces aligned with these variables shift model outputs. The findings suggest that LMs compare quantities via a bag of heuristics operating over numerals and units, rather than first converting both expressions into a precise shared-scale representation. Authors: Mutsumi Sasaki, Go Kamoda, Ryosuke Takahashi, Kosuke Sato, Kentaro Inui, Keisuke Sakaguchi, Benjamin Heinzerling.

Paper Overview

  • Field: NLP
  • Authors: Mutsumi Sasaki, Go Kamoda, Ryosuke Takahashi, Kosuke Sato, Kentaro Inui, Keisuke Sakaguchi, Benjamin Heinzerling
  • Posted: 2026-06-02
  • arXiv: 2606.03982
  • Abstract

    Quantities with measurement units, such as 110 cm and 1.2 m, require language models (LMs) to combine a numeral with a symbolic unit scale. Here, we study how LMs compare such quantities in controlled settings spanning several unit systems.

    Key Findings

  • Accuracy degrades near the comparison boundary, where small changes in value determine the correct answer.
  • The resulting errors are systematic: linear surrogate models predict LM preferences from numerical-difference and unit-scale-difference cues.
  • Causal interventions on subspaces aligned with these variables shift the models' outputs.
  • The results suggest that LMs compare quantities through a bag of heuristics over numerals and units, rather than first converting both expressions to an exact shared-scale representation.
---

*Auto-collected on 2026-06-04.*

Tags

#language-models#nlp#quantitative-reasoning#unit-conversion#interpretability#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177980808