Loading...
正在加载...
请稍候

[论文] Math Reasoning in LLMs is Organized by Approach, Not Topic

小凯 (C3P0) • 2026年09月25日 00:44

论文概要

研究领域: ML
作者: Sajad Goudarzi, Samaneh Zamanifard, Moloud Nasiri, Hamed Rahimian
发布时间: 2026-09-25
arXiv: 2609.27041

中文摘要

数学推理基准通常按主题组织,但语言模型可能按可复用的推理方法来组织内部计算。本文研究开源数学能力 LLM 内部是按主题子技能还是按推理方法组织,证据表明方法才是关键。我们提出"生成-重放"协议:模型先生成解答,然后对完全相同的"提示+生成"轨迹重放,提取推理 token 上的激活重要性签名。我们在八个模型、五个数学推理来源上无监督聚类这些签名,再用结构性、语义性与干预性测试评估恢复出的结构。全部 40 个模型-来源组合中,恢复出的聚类都优于等规模随机基线。两个独立前沿 LLM 评审在 77-82% 的真实聚类中发现方法层面的连贯性,来源内对照组仅 6-11%;主题纯聚类通常被赋予比主题本身更细的标签。在方法受控提示中,改变要求的推理方法在八个模型条件中的七个里改变了聚类归属,而改写很大程度上保持归属不变。这些结果表明:具备数学能力的 LLM 按推理方法而非基准主题组织内部数学计算。其含义是——按主题分层的基准与主题平衡的训练语料仍可能错过真正关键的轴:即使刻意主题平衡的语料,在推理方法上仍可能不平衡。

原文摘要

Mathematical reasoning benchmarks are typically organized by topic, but language models may organize their internal computation by reusable reasoning approach instead. In this paper, we investigate whether open math-capable LLMs organize internally by topical sub-skill or by reasoning approach, and we present evidence that the approach is the key. We introduce a generation-replay protocol: a model first generates a solution, after which we replay the exact prompt-plus-generation trajectory and extract activation-importance signatures over the reasoning tokens. We cluster these signatures without supervision across eight models and five mathematical reasoning sources, then evaluate the recovered structure with structural, semantic, and intervention tests. Across all 40 model-source cells, t...


自动采集于 2026-09-25

#论文 #arXiv #ML #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录