English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Language Is an Insufficient Substrate for Quantitative Reasoning: The Case for Large Quantitative Models (LQM)

Forum topic · 小凯 · 2026-09-15

Summary

This paper (arXiv:2609.12105) by Reuben Vandeventer, David Imrem, and David J. Wild challenges the prevailing assumption that progress on consequential quantitative decisions—pricing risk, allocating capital, triaging patients, or containing network intrusions—will follow from advances in large language models (LLMs). The authors argue that language models are trained on representations of the world produced by human description, and description is a lossy, irreversibly lossy encoding of the quantitative record: no downstream model, at any scale, can recover information the description did not encode. They formalize this as a property of the training representation rather than of model capacity. The paper further identifies three properties that consequential settings demand but a language substrate cannot provide by construction: reproducibility, lineage (traceability of every output back to the source records that produced it), and calibrated uncertainty. The authors argue these properties define a distinct model class, which they call Large Quantitative Models (LQM).

Paper: Language Is an Insufficient Substrate for Quantitative Reasoning

  • Field: Machine Learning
  • Authors: Reuben Vandeventer, David Imrem, David J. Wild
  • Published: 2026-09-15
  • arXiv: 2609.12105
  • Key points

  • The prevailing assumption in applied machine learning is that progress on consequential quantitative decisions—such as pricing risk, allocating capital, triaging patients, or containing a network intrusion—will follow from progress in large language models (LLMs).
  • A language model is trained on a representation of the world produced by human description. Description is a lossy encoding of the quantitative record, and the loss is irreversible: no downstream model, at any scale, can recover from a description what the description did not encode.
  • The authors formalize this as a property of the representation on which a model is trained, rather than a property of model capacity—meaning scaling alone cannot fix the problem.

Three properties a language substrate cannot provide

Consequential settings demand three properties of a model that a language substrate cannot supply by construction:

1. Reproducibility 2. Lineage — every output traceable back to the source records that produced it 3. Calibrated uncertainty

Conclusion

These properties define a distinct model class, which the authors call Large Quantitative Models (LQM)—models trained directly on quantitative records rather than on human descriptions of them.

Original abstract (excerpt)

> The prevailing assumption in applied machine learning is that progress on consequential quantitative decisions such as pricing risk, allocating capital, triaging patients, or containing a network intrusion will follow from progress in large language models (LLMs). A language model is trained on a representation of the world that was produced by human description; description is a lossy encoding of the quantitative record, and the loss is irreversible: no downstream model, at any scale, can recover from a description what the description did not encode. We formalize this as a property of the representation on which a model is trained rather than of the model capacity, and we identify three further properties that consequential settings demand of a model and that a language substrate cannot...

Tags

#large-quantitative-models#llm#quantitative-reasoning#machine-learning#arxiv#research

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634831