[论文] Language Is an Insufficient Substrate for Quantitative Reasoning, and ...
研究领域: ML 作者: Reuben Vandeventer, David Imrem, David J. Wild 发布时间: 2026-09-15 arXiv: 2609.12105
论文概要
研究领域: ML 作者: Reuben Vandeventer, David Imrem, David J. Wild 发布时间: 2026-09-15 arXiv: 2609.12105
中文摘要
应用机器学习领域的主流假设是:定价风险、资本配置、患者分诊或遏制网络入侵等重大量化决策的进展将跟随大型语言模型(LLM)的进展而来。然而语言模型所训练的世界表征由人类描述产生;描述是对量化记录的有损编码,且这种损失不可逆——任何规模的下游模型都无法从描述中恢复描述未编码的信息。我们将此形式化为模型所训练表征的一种属性,而非模型容量的属性,并进一步指出重大场景所要求、而语言基底在构造上无法提供的三个属性:可复现性、每个输出回溯至产生它的源记录的血缘关系,以及校准的不确定性。我们认为这些属性定义了一个独特的模型类别,称之为大型量化模型(LQM)。
原文摘要
The prevailing assumption in applied machine learning is that progress on consequential quantitative decisions such as pricing risk, allocating capital, triaging patients, or containing a network intrusion will follow from progress in large language models (LLMs). A language model is trained on a representation of the world that was produced by human description; description is a lossy encoding of the quantitative record, and the loss is irreversible: no downstream model, at any scale, can recover from a description what the description did not encode. We formalize this as a property of the representation on which a model is trained rather than of the model capacity, and we identify three further properties that consequential settings demand of a model and that a language substrate cannot ...
*自动采集于 2026-09-15*
#论文 #arXiv #ML #小凯